What Happened to "GLM 5.2 Air"
During a July 2026 data incident, the LLMCheck catalog listed a "GLM 5.2 Air" described as a Mac-runnable distillation of Zhipu's flagship. When the catalog was reconciled against official sources in August 2026, no such release could be verified: Zhipu AI has never shipped an Air variant of the 5.x series. The specs attached to that entry — 106B total parameters, 12B active, MIT license, ~30 tok/s on a 64 GB Mac (estimated) — match a model that very much does exist: GLM-4.5-Air.
The rest of this page covers the verified record. If you came here to run a big GLM on Apple Silicon, the good news is that almost everything practical still holds — you were just one version number off.
What GLM-4.5-Air Actually Is
GLM-4.5-Air is the compact sibling of GLM-4.5 (355B-A32B), Zhipu AI's agent-focused model family. The Air configuration is 106B total parameters with only 12B active at any given step. That MoE design is the whole trick: the model holds 106B parameters' worth of knowledge in memory, but each generated token only routes through 12B of them. You pay the RAM cost of a large model while paying the compute cost of a small one — which is precisely what makes it viable on a Mac.
It ships with a 128K-token context window and the MIT license, the most permissive option in the open-weights world. The 4.5 family was trained with agentic tool use in mind — planning, function calling, multi-step coding loops — and Air inherits that orientation. According to the LLMCheck index, it remains the deepest-reasoning GLM you can fit into 64 GB of unified memory.
The key numbers: 106B total / 12B active, MIT license, 128K context. Everything else on this page — RAM tiers, speeds, install commands — follows from those three facts.
Hardware Requirements
At 4-bit quantization, GLM-4.5-Air's weights occupy roughly 60 GB of unified memory. Add the OS footprint, your apps, and a working context buffer, and 64 GB is the absolute floor. Here is how the tiers shake out:
| RAM Tier | Verdict | What You Get |
|---|---|---|
| 32 GB or less | Not supported | Weights do not fit at any usable quant — run GLM-4.7-Flash instead. |
| 64 GB | Minimum | 4-bit runs, tightly; keep context modest (~16–32K) and quit heavy apps. |
| 96–128 GB | Comfortable | Long contexts, or a higher-fidelity quant, without memory pressure. |
| 192 GB (Ultra) | Ideal | Highest throughput, full 128K context, room for parallel models. |
A practical note for 64 GB owners: the model fits, but unified memory is shared with everything else on the system. Cap your context length to avoid swapping, and expect to nudge macOS's GPU wired-memory limit for the largest builds. The M5 Max and M4 Ultra with 128 GB+ are where GLM-4.5-Air feels effortless.
Estimated Speeds by Apple Silicon Chip
Because only 12B parameters are active per token, GLM-4.5-Air punches far above what a dense 106B model could ever manage on a Mac. The figures below are LLMCheck index estimates from our published model — not lab measurements (see how we estimate):
| Chip / RAM | Runtime | Throughput (est.) |
|---|---|---|
| M4 Max 128 GB | MLX | ~26 tok/s |
| M5 Max 64 GB | Ollama | ~30 tok/s |
| M5 Max 128 GB | MLX | ~34 tok/s |
| M4 Ultra 192 GB | MLX | ~38 tok/s |
All four configurations sit comfortably above conversational reading speed (~10–15 tok/s). Even the 64 GB M5 Max — the cheapest viable machine — holds an estimated 30 tok/s, which is enough for interactive coding and chat. Chip-by-chip detail lives on the model pages: GLM-4.5-Air on M5 Max, on M4 Max, and on M4 Ultra.
Step-by-Step Install (Ollama + MLX)
There are two clean paths to running GLM-4.5-Air locally. Ollama is the fastest to set up; MLX squeezes out a few extra tokens per second on Apple Silicon.
Option A — Ollama (easiest)
If you do not already have Ollama, install it from the software page or via Homebrew, then pull and run the model in a single command:
That single ollama run command downloads the 4-bit weights, loads them into unified memory, and drops you into an interactive prompt. To serve it to other apps over the local API, run ollama serve and point your client at http://localhost:11434.
Option B — MLX (fastest on Apple Silicon)
MLX is Apple's own array framework, tuned for the unified-memory architecture. It typically delivers a few extra tok/s at the same quantization. Install mlx-lm and pull the community-quantized build:
For an interactive chat loop or an OpenAI-compatible server, use mlx_lm.chat or mlx_lm.server respectively. The first invocation downloads the weights from Hugging Face; subsequent runs load from the local cache.
On a 64 GB Mac, start with Ollama at 4-bit — it is the safest, most predictable path. Move to MLX once you want maximum throughput and are comfortable managing Python environments.
Recommended Quantization
Quantization is the lever that trades memory and speed against output fidelity. For GLM-4.5-Air on a Mac, two settings cover almost every case:
- 4-bit (Q4_K_M / MLX 4-bit) — roughly 60 GB of weights. The standard build shipped by the community and the only realistic choice on a 64 GB Mac. Quality loss versus full precision is small and rarely noticeable in coding or chat.
- 3-bit — roughly 48 GB. Worth considering on exactly 64 GB if you need more context headroom; expect some fidelity loss on the hardest multi-step reasoning.
- 5–6-bit — 128 GB+ machines only, where the extra headroom is free. It tightens up reasoning on hard problems and reduces occasional formatting drift in long agentic runs.
Skip 8-bit and full precision on a single Mac — the memory cost is not justified by the marginal quality gain for this model.
The Newer Mac GLM: GLM-4.7-Flash
Since GLM-4.5-Air shipped, Zhipu has released a newer Mac-friendly option: GLM-4.7-Flash, a 31B dense model under the same MIT license. At 4-bit it needs roughly 18 GB of RAM, which puts it within reach of 24–32 GB Macs — a much larger audience than Air's 64 GB floor. According to the LLMCheck index, it runs an estimated ~26 tok/s on an M5 Max (see GLM-4.7-Flash on M5 Max).
| Factor | GLM-4.5-Air | GLM-4.7-Flash |
|---|---|---|
| Total / Active Params | 106B / 12B (MoE) | 31B dense |
| Min RAM (4-bit) | 64 GB | 24–32 GB |
| Reasoning Depth | Deeper (106B of knowledge) | Strong, lighter |
| Speed on M5 Max (est.) | ~30–34 tok/s | ~35 tok/s |
| License | MIT | MIT |
Pick GLM-4.5-Air if…
You have 64 GB+ of RAM and want the deepest GLM reasoning available locally — hard agentic coding, whole-repo analysis, or long-document work where the 106B knowledge pool matters more than raw accessibility.
Pick GLM-4.7-Flash if…
You are on a 24–32 GB Mac, or you want the newest GLM without the memory bill. It handles everyday coding and chat well and leaves your machine free to do other work at the same time.
What About the Full GLM 5.2?
The real GLM 5.2 — Zhipu's 753B flagship, released in June 2026 — is a server-class system, and there is no Air variant of it. That said, the community has pushed it onto the very largest Macs: a 1-bit GGUF reaches ~22 tok/s on a 256 GB M3 Ultra and a 4-bit MLX build ~15 tok/s on a 512 GB M3 Ultra (both community-reported; see GLM 5.2 on M3 Ultra). For anything below that tier, GLM-4.5-Air is the flagship-adjacent GLM you can actually run — and the full comparison of GLM 5.2 against DeepSeek's V4 line lives in our open-frontier showdown.
If Air is the GLM you end up running, three caveats are worth knowing before you commit:
- 64 GB is genuinely the floor for Air. The model fits, but it leaves little room for large contexts or other heavy apps. Memory pressure and swapping will tank your throughput if you are not disciplined.
- The 128K context is RAM-bound. Filling the full window costs unified memory on top of the weights. Only 128 GB+ machines can use it without compromise.
- MoE routing adds variance. Because different tokens activate different experts, throughput can fluctuate slightly run to run more than a dense model would.
None of these are dealbreakers. For a model that brings 106B parameters of MIT-licensed reasoning to a single Mac, GLM-4.5-Air remains the most capable local GLM in its RAM class as of August 2026, according to the LLMCheck index.