What Happened to "GLM 5.2 Air"

During a July 2026 data incident, the LLMCheck catalog listed a "GLM 5.2 Air" described as a Mac-runnable distillation of Zhipu's flagship. When the catalog was reconciled against official sources in August 2026, no such release could be verified: Zhipu AI has never shipped an Air variant of the 5.x series. The specs attached to that entry — 106B total parameters, 12B active, MIT license, ~30 tok/s on a 64 GB Mac (estimated) — match a model that very much does exist: GLM-4.5-Air.

The rest of this page covers the verified record. If you came here to run a big GLM on Apple Silicon, the good news is that almost everything practical still holds — you were just one version number off.

What GLM-4.5-Air Actually Is

GLM-4.5-Air is the compact sibling of GLM-4.5 (355B-A32B), Zhipu AI's agent-focused model family. The Air configuration is 106B total parameters with only 12B active at any given step. That MoE design is the whole trick: the model holds 106B parameters' worth of knowledge in memory, but each generated token only routes through 12B of them. You pay the RAM cost of a large model while paying the compute cost of a small one — which is precisely what makes it viable on a Mac.

It ships with a 128K-token context window and the MIT license, the most permissive option in the open-weights world. The 4.5 family was trained with agentic tool use in mind — planning, function calling, multi-step coding loops — and Air inherits that orientation. According to the LLMCheck index, it remains the deepest-reasoning GLM you can fit into 64 GB of unified memory.

The key numbers: 106B total / 12B active, MIT license, 128K context. Everything else on this page — RAM tiers, speeds, install commands — follows from those three facts.

Hardware Requirements

At 4-bit quantization, GLM-4.5-Air's weights occupy roughly 60 GB of unified memory. Add the OS footprint, your apps, and a working context buffer, and 64 GB is the absolute floor. Here is how the tiers shake out:

RAM Tier Verdict What You Get
32 GB or less Not supported Weights do not fit at any usable quant — run GLM-4.7-Flash instead.
64 GB Minimum 4-bit runs, tightly; keep context modest (~16–32K) and quit heavy apps.
96–128 GB Comfortable Long contexts, or a higher-fidelity quant, without memory pressure.
192 GB (Ultra) Ideal Highest throughput, full 128K context, room for parallel models.

A practical note for 64 GB owners: the model fits, but unified memory is shared with everything else on the system. Cap your context length to avoid swapping, and expect to nudge macOS's GPU wired-memory limit for the largest builds. The M5 Max and M4 Ultra with 128 GB+ are where GLM-4.5-Air feels effortless.

Estimated Speeds by Apple Silicon Chip

Because only 12B parameters are active per token, GLM-4.5-Air punches far above what a dense 106B model could ever manage on a Mac. The figures below are LLMCheck index estimates from our published model — not lab measurements (see how we estimate):

Chip / RAM Runtime Throughput (est.)
M4 Max 128 GB MLX ~26 tok/s
M5 Max 64 GB Ollama ~30 tok/s
M5 Max 128 GB MLX ~34 tok/s
M4 Ultra 192 GB MLX ~38 tok/s

All four configurations sit comfortably above conversational reading speed (~10–15 tok/s). Even the 64 GB M5 Max — the cheapest viable machine — holds an estimated 30 tok/s, which is enough for interactive coding and chat. Chip-by-chip detail lives on the model pages: GLM-4.5-Air on M5 Max, on M4 Max, and on M4 Ultra.

Step-by-Step Install (Ollama + MLX)

There are two clean paths to running GLM-4.5-Air locally. Ollama is the fastest to set up; MLX squeezes out a few extra tokens per second on Apple Silicon.

Option A — Ollama (easiest)

If you do not already have Ollama, install it from the software page or via Homebrew, then pull and run the model in a single command:

# Install Ollama (skip if already installed) $ brew install ollama # Pull and run GLM-4.5-Air — first run downloads ~60 GB $ ollama run glm-4.5-air

That single ollama run command downloads the 4-bit weights, loads them into unified memory, and drops you into an interactive prompt. To serve it to other apps over the local API, run ollama serve and point your client at http://localhost:11434.

Option B — MLX (fastest on Apple Silicon)

MLX is Apple's own array framework, tuned for the unified-memory architecture. It typically delivers a few extra tok/s at the same quantization. Install mlx-lm and pull the community-quantized build:

# Install the MLX language-model toolkit $ pip install mlx-lm # Run GLM-4.5-Air from the MLX community repo $ mlx_lm.generate \ --model mlx-community/GLM-4.5-Air-4bit \ --prompt "Refactor this function for readability:" \ --max-tokens 1024

For an interactive chat loop or an OpenAI-compatible server, use mlx_lm.chat or mlx_lm.server respectively. The first invocation downloads the weights from Hugging Face; subsequent runs load from the local cache.

On a 64 GB Mac, start with Ollama at 4-bit — it is the safest, most predictable path. Move to MLX once you want maximum throughput and are comfortable managing Python environments.

Recommended Quantization

Quantization is the lever that trades memory and speed against output fidelity. For GLM-4.5-Air on a Mac, two settings cover almost every case:

Skip 8-bit and full precision on a single Mac — the memory cost is not justified by the marginal quality gain for this model.

The Newer Mac GLM: GLM-4.7-Flash

Since GLM-4.5-Air shipped, Zhipu has released a newer Mac-friendly option: GLM-4.7-Flash, a 31B dense model under the same MIT license. At 4-bit it needs roughly 18 GB of RAM, which puts it within reach of 24–32 GB Macs — a much larger audience than Air's 64 GB floor. According to the LLMCheck index, it runs an estimated ~26 tok/s on an M5 Max (see GLM-4.7-Flash on M5 Max).

Factor GLM-4.5-Air GLM-4.7-Flash
Total / Active Params 106B / 12B (MoE) 31B dense
Min RAM (4-bit) 64 GB 24–32 GB
Reasoning Depth Deeper (106B of knowledge) Strong, lighter
Speed on M5 Max (est.) ~30–34 tok/s ~35 tok/s
License MIT MIT

Pick GLM-4.5-Air if…

You have 64 GB+ of RAM and want the deepest GLM reasoning available locally — hard agentic coding, whole-repo analysis, or long-document work where the 106B knowledge pool matters more than raw accessibility.

Pick GLM-4.7-Flash if…

You are on a 24–32 GB Mac, or you want the newest GLM without the memory bill. It handles everyday coding and chat well and leaves your machine free to do other work at the same time.

What About the Full GLM 5.2?

The real GLM 5.2 — Zhipu's 753B flagship, released in June 2026 — is a server-class system, and there is no Air variant of it. That said, the community has pushed it onto the very largest Macs: a 1-bit GGUF reaches ~22 tok/s on a 256 GB M3 Ultra and a 4-bit MLX build ~15 tok/s on a 512 GB M3 Ultra (both community-reported; see GLM 5.2 on M3 Ultra). For anything below that tier, GLM-4.5-Air is the flagship-adjacent GLM you can actually run — and the full comparison of GLM 5.2 against DeepSeek's V4 line lives in our open-frontier showdown.

If Air is the GLM you end up running, three caveats are worth knowing before you commit:

None of these are dealbreakers. For a model that brings 106B parameters of MIT-licensed reasoning to a single Mac, GLM-4.5-Air remains the most capable local GLM in its RAM class as of August 2026, according to the LLMCheck index.