Why this is still a landmark

Open-weights models have beaten closed models on individual benchmarks before. Qwen has topped MMLU. DeepSeek has led AIME math. But none of those wins touched the benchmark that best predicts whether a model can do an engineer’s job: resolving real, multi-file GitHub issues autonomously. In June 2026, GLM 5.2 — released by Zhipu AI, the Beijing lab also known as Z.ai — became the first open-weights model to report a SWE-Bench Pro score above GPT-5 and Claude Opus 4.6.

The August benchmark-integrity debate added the necessary asterisk. Zhipu’s 68.5% was produced on Zhipu’s own agent scaffold, and agentic benchmark scores are not comparable across scaffolds. On Scale’s standardized SEAL harness — which runs every model through the same agent loop — GLM 5.2 scores about 62%: the top open-weights result on record, but behind the closed frontier when everyone runs the same code. The claim “first open model to beat GPT-5 and Claude” is scaffold-dependent, and this page now says so plainly.

The claim, stated precisely: on Zhipu’s own scaffold, GLM 5.2 reports 68.5% on SWE-Bench Pro — above GPT-5’s and Claude Opus 4.6’s vendor-reported numbers. On Scale’s standardized SEAL harness it scores ~62%: the best open-weights result, behind the closed frontier. GLM 5.2 leads open weights either way; whether it “beats GPT-5” depends entirely on whose harness you trust.

What survives the correction still matters. A model that is simultaneously (a) the strongest open-weights coding model on a standardized harness, (b) fully open under MIT with no usage caps, and (c) — as of August 2026 — runnable on the largest Apple Silicon Macs via community quantizations is a first. The open-closed gap on real engineering work, which ran 5–10 points a year ago, is now a few points on a common harness. That narrowing, not any single headline number, is the landmark.


The headline benchmark: 68.5% — on whose harness?

SWE-Bench Pro is the hardened successor to SWE-Bench Verified. Where Verified measured single-patch correctness on a curated set, Pro tests the full agentic loop — clone the repository, read the issue, locate the relevant files, edit across multiple modules, run the test suite, read the failures, and iterate until the patch passes. That loop is exactly why its scores are scaffold-sensitive: the agent framework, tool set, and retry budget wrapped around the model can move the number by several points. Here is the state of reported scores as of August 2026, with provenance attached:

SWE-Bench Pro — reported scores, August 2026. Scores from different scaffolds are NOT directly comparable; the SEAL row is the standardized reference.
Model Type SWE-Bench Pro Harness
GLM 5.2Open (MIT)68.5%Zhipu’s own scaffold (self-reported)
Qwen3.8-2.4T-A95BOpen (custom)67.7%Vendor-reported
GPT-5Closed67.0%Vendor-reported
Claude Opus 4.6Closed65.1%Vendor-reported
GLM 5.2Open (MIT)~62%Scale SEAL (standardized) — top open score
Laguna S 2.1Open (OpenMDW)59.4%Public set, vendor-reported
Muse Glimmer 30BOpen (Apache 2.0)51.2%Vendor-reported

The number is real, but scaffold-dependent. Nobody credibly disputes that Zhipu’s harness produced 68.5%. What the August benchmark-integrity debate established is that vendor scaffolds are tuned to their own models — different tool APIs, different retry budgets, different context management — so cross-vendor comparisons of self-reported agentic scores are close to meaningless. Scale’s SEAL leaderboard exists to fix exactly this: one harness, every model. On SEAL, GLM 5.2’s ~62% leads every open-weights model but sits below the closed frontier’s results on the same harness.

What survives the correction: GLM 5.2 is the top open-weights coding model on the standardized harness, and the open-closed gap on SEAL is a few points — the smallest it has ever been. The historical pattern was open models trailing by 5–10 points on agentic coding. That era is over, even if the “open beats closed” headline was premature.

Why coding is the bellwether: SWE-Bench Pro requires planning, tool use, long-context code comprehension, and error recovery in a single loop — the same capabilities that drive agentic performance across domains. That is also precisely why the harness matters so much: the loop is half the score. When you read any agentic benchmark claim in 2026, the first question is no longer “what did it score?” but “on whose scaffold?”


Architecture deep-dive: 753B/40B MoE

GLM 5.2 is a sparse mixture-of-experts model: 753 billion total parameters, of which only about 40 billion activate per token. That ratio — roughly 5.3% of the network firing on any given forward pass — is what makes a model this large tractable to serve at all. The dense-equivalent quality comes from the full 753B parameter pool; the inference cost tracks the 40B active slice.

Architecture

Total params753B
Active params40B (5.3%)
TypeSparse MoE
ReasoningHybrid mode
Context256K
Training tokens22T

Provenance

LabZhipu AI (Z.ai)
ReleasedJune 2026
LicenseMIT
AcceleratorHuawei Ascend
WeightsHugging Face
Commercial useUnrestricted

Hybrid reasoning is the capability unlock. Like the strongest 2026 models, GLM 5.2 ships with a hybrid reasoning mode that lets it decide, per query, whether to answer directly or to spend tokens on an explicit chain of thought before responding. On easy queries it stays fast; on SWE-Bench-style multi-step problems it switches into extended reasoning automatically. This is the mechanism behind the agentic-coding score — the model plans before it edits, then reflects on test failures before it retries.

The training story is geopolitically notable. GLM 5.2 was trained on 22 trillion tokens on a domestic accelerator cluster built on Huawei Ascend silicon — not Nvidia. This matters beyond benchmarks: it demonstrates that a frontier-class model can now be trained end-to-end outside the Nvidia ecosystem. For the open-weights movement, it means the supply chain for state-of-the-art models is diversifying, which makes the open frontier more resilient to any single hardware bottleneck.

256K context is workhorse-sized, not record-setting. A 256K context window comfortably holds a large codebase, a long document set, or an extended agentic trajectory. It is not the 1M-token frontier that some hosted models chase, but for the coding-agent use case GLM 5.2 is built for, 256K is more than enough to hold a repository’s relevant files plus a full reasoning trace.


Full benchmark suite vs GPT-5 & Claude Opus 4.6

SWE-Bench Pro is the headline, but it is not the whole picture. Here is GLM 5.2 across the standard 2026 benchmark suite against the two closed models it most directly challenges. Static evals (MMLU, HumanEval, GPQA, AIME) are far less scaffold-sensitive than agentic ones, so these comparisons are on firmer ground — though all closed-model figures are vendor-reported. The pattern: GLM 5.2 leads or ties on code generation and competition math, and trades within a point or two everywhere else.

GLM 5.2 vs GPT-5 vs Claude Opus 4.6 — August 2026. Green = best in row. *Zhipu’s own scaffold; on Scale’s standardized SEAL harness GLM 5.2 scores ~62%.
Benchmark GLM 5.2 GPT-5 Claude Opus 4.6
SWE-Bench Pro68.5%*67.0%65.1%
SWE-Verified84%83%83%
MMLU92%92%91%
HumanEval97%96%96%
AIME 202694%93%90%
GPQA-Diamond89%89%88%
Arena ELO~1565~1580~1572
LicenseMIT (open)ClosedClosed

Where GLM 5.2 wins clearly: code generation (HumanEval 97%) and competition math (AIME 94%) — static evals where the harness is not a confound. On agentic coding it is the top open-weights model on the standardized SEAL harness, which is a first, even though the closed frontier stays ahead there.

Where it ties: MMLU (92%), GPQA-Diamond (89%), and SWE-Verified (84%) are statistical dead heats with the closed frontier — the differences are inside the margin of error of a single evaluation run.

Where it trails: Arena ELO. GLM 5.2’s ~1565 sits just behind GPT-5 (~1580) and Claude Opus 4.6 (~1572). Arena ELO measures human-preferred general chat — tone, helpfulness, formatting, refusal calibration — and the closed labs still hold a small edge there. If your use case is consumer chat, the closed models remain marginally preferred; if your use case is code and reasoning, GLM 5.2 is at or near the frontier — and it is the model you can download.


The license story — MIT vs the closed frontier

The benchmark numbers get the headlines, but the MIT license is what actually changes the industry. GPT-5 and Claude Opus 4.6 are accessible only through paid APIs. You rent intelligence by the token, you cannot inspect the weights, you cannot fine-tune on your private data without sending it to a vendor, and your unit economics are permanently tied to someone else’s pricing.

GLM 5.2 inverts all of that. MIT is the most permissive mainstream license in existence. It grants:

Compare this to the “open-ish” licenses that dominated 2025 — MAU caps, prohibited-use lists, no-compete clauses — or even to 2026’s new revenue-gated licenses like Kimi K3’s and Qwen3.8’s. GLM 5.2 has none of that. It holds the top open-weights score on the standardized SWE-Bench Pro harness, and it is also one of the most permissively licensed frontier models. For a startup deciding whether to build on a closed API or a self-hosted open model, the open option is no longer a downgrade — it is a scaffold-tuning project away from parity.

The startup math: A coding-agent product that would cost $40,000/month in GPT-5 API fees at scale can run on rented or owned GLM 5.2 inference for a fraction of that — with full data control and the freedom to fine-tune. MIT licensing turns the frontier from an operating expense into a capital decision.


Can you run GLM 5.2 on a Mac?

When this review first ran, the honest answer was “no — server-class only.” As of August 2026, the answer has changed at the extremes. At 4-bit quantization the full 753B model still occupies roughly 400 GB, which remains 8×H100 territory for production serving. But the community quantization scene has pushed the full model onto the two largest Apple Silicon configurations:

Full GLM 5.2 (753B) on Apple Silicon — community-reported, August 2026.
Mac Unified RAM Quant / Runtime Speed
M3 Ultra256 GB1-bit GGUF (llama.cpp)~22 tok/s
M3 Ultra512 GB4-bit MLX~15 tok/s

Both figures are community-reported. Note the inversion: the 1-bit build on the 256 GB machine is faster than the 4-bit build on the 512 GB machine, because it moves far less memory per token — but 1-bit quality is materially degraded, and neither build is the benchmark-grade model that scored the numbers above. Treat these as remarkable proofs of capability for experimentation and offline agent work, not as a substitute for hosted inference. Full details are on the GLM 5.2 on M3 Ultra model page.

For every other Mac, the answer is the smaller GLM models. Zhipu’s Mac-runnable line is GLM-4.5-Air — a 106B-A12B MoE that fits a 64 GB machine — and the newer GLM-4.7-Flash, a 31B dense model for the 32 GB tier. Both are MIT licensed and covered in the next section.

⚡ Run it at full size — rent a GPU

Full GLM 5.2 (753B-A40B) needs ~400 GB at 4-bit — beyond every Mac except a 512 GB M3 Ultra, and beyond benchmark-grade quality at 1-bit. To run the full, unquantized model, rent a datacenter GPU by the minute on Vast.ai — often 5–6× cheaper than AWS or GCP, with H100s and B200s available on demand.

Vast.ai referral link — we may earn a small commission at no extra cost to you. It never influences our reviews or rankings.


GLM-4.5-Air and GLM-4.7-Flash — the practical Mac answers

For the overwhelming majority of LLMCheck readers, the model you will actually run is not the 753B flagship — it is one of the two Mac-sized models in the GLM line. A correction first: an earlier version of this section described a “GLM 5.2 Air.” No 5.x Air model exists; the real 106B-A12B Mac-runnable model is GLM-4.5-Air, and the LLMCheck index has been corrected accordingly.

GLM-4.5-Air

Total params106B
Active params12B
TypeSparse MoE
LicenseMIT
Min RAM64 GB
Speed (64 GB Mac)~30 tok/s est.

GLM-4.7-Flash

Total params31B
TypeDense
LicenseMIT
Size @ Q4~18 GB
Min RAM32 GB
StatusNewest Mac GLM

GLM-4.5-Air is the 64 GB-tier pick. With only 12B parameters active per token, it runs far quicker than a 106B dense model would — according to the LLMCheck index, an estimated ~30 tok/s on a 64 GB Apple Silicon Mac at 4-bit. That is comfortably interactive for a coding agent that reads files, proposes edits, and reacts to test output, and it remains one of the strongest models that fits consumer-accessible hardware. See the GLM-4.5-Air on M5 Max page for per-chip figures.

GLM-4.7-Flash is the newer, smaller option. A 31B dense model at roughly 18 GB in 4-bit, it fits the 32 GB Mac tier that GLM-4.5-Air cannot reach. Zhipu has not published M-series throughput figures; the GLM-4.7-Flash on M5 Max page carries the LLMCheck estimates.

Installation on a Mac:

# MLX (Apple Silicon, 64 GB+ for Air)
pip install mlx-lm
mlx_lm.generate --model mlx-community/GLM-4.5-Air-4bit \
--prompt "Read this repo and add a rate-limiter to the API layer"

# LM Studio: search "GLM-4.5-Air" or "GLM-4.7-Flash" in the Discover tab

For a developer who wants a self-hosted coding assistant that runs entirely on their own Mac, costs nothing per token, and carries no licensing risk, GLM-4.5-Air remains the default GLM-line recommendation as of August 2026 — with GLM-4.7-Flash as the pick for 32 GB machines.


GLM 5.2 vs the open field

GLM 5.2 did not arrive into an empty field, and the two months since its release have been the busiest of the year for open weights. Positioning it against the August 2026 field clarifies exactly what it is for.

vs DeepSeek V4-Flash. The practical rival for Mac users. V4-Flash (July 2026, MIT, 284B-A13B) posts a self-reported 82.7 on Terminal Bench 2.1 and sits in the top open tier on the Artificial Analysis index — and its 2-bit MLX build (96.5 GB) runs on a 128 GB Mac, where community reports put it around 39 tok/s on an M5 Max. GLM 5.2 holds the higher standardized SWE-Bench Pro score; V4-Flash runs on hardware people actually own. See DeepSeek V4-Flash on M5 Max.

vs Kimi K3. Moonshot’s 2.8-trillion-parameter MoE (July 2026) is the largest open-weight model ever published — roughly 1.56 TB of weights — and ranks #1 among open models on LiveBench Coding. But its custom “Kimi K3 License” carries a revenue gate for large MaaS providers, and at that size it is not Mac-runnable at any quantization. GLM 5.2 is the more permissive and more deployable of the two.

vs Qwen3.8-2.4T-A95B. Alibaba’s open-weights flagship (August 2026) vendor-reports 67.7 on SWE-Bench Pro and 92.6 on GPQA — but it arrived under a custom revenue-threshold license, the first non-Apache Qwen release, and its smallest quantization is 397 GB, out of Mac reach. Against it, GLM 5.2 offers a cleaner license and a smaller footprint at comparable reported capability.

vs Inkling-Small. Thinking Machines Lab’s 276B-A12B model (July 2026, Apache 2.0) holds the open-weight record on SWE-bench Verified at 80.2% — a harness-cleaner benchmark than Pro — and its 2-bit MLX build (88.4 GB) fits a 128 GB Mac Studio, though it is MLX-only for now. For agentic coding on a Mac you can buy today, Inkling-Small is arguably the sharper pick; GLM 5.2’s advantage is the full-strength flagship being self-hostable under MIT.


How to deploy GLM 5.2

There are three realistic deployment paths depending on what hardware you have and how much quality you need.

Path 1 — Full GLM 5.2 on your own server

For the 753B flagship at production quality you need a GPU node with roughly 400 GB+ of memory at 4-bit — an 8×H100 (640 GB) or 8×H200 box. Serve it with vLLM or SGLang for production throughput.

# vLLM (recommended for production serving)
pip install vllm
vllm serve zai-org/GLM-5.2 \
--tensor-parallel-size 8 \
--max-model-len 262144

Path 2 — Full GLM 5.2 on an M3 Ultra (community quants)

On a 256 GB M3 Ultra, the community 1-bit GGUF runs under llama.cpp at ~22 tok/s (community-reported); on a 512 GB M3 Ultra, the 4-bit MLX build reaches ~15 tok/s (community-reported). Expect real quality loss at 1-bit — these builds are for experimentation and offline agent work, not benchmark-grade output. For every other Mac, run GLM-4.5-Air (64 GB+) or GLM-4.7-Flash (32 GB+) instead — see the section above.

# M3 Ultra 256 GB — community 1-bit GGUF via llama.cpp
llama-server -m glm-5.2-1bit.gguf -c 32768

# M3 Ultra 512 GB — community 4-bit MLX build
mlx_lm.generate --model <community-GLM-5.2-4bit-repo> --prompt "..."

Path 3 — Rent the full model in the cloud

If you want full GLM 5.2 quality without owning a cluster, rent an 8×H100 node by the hour from any major GPU cloud. As of August 2026 that runs in the low single digits of dollars per H100-hour, so a full 8-GPU node is on the order of $15–25/hour — cost-effective for batch jobs, evaluation runs, or bursty agent workloads, and you tear it down when you are done. Because the weights are MIT-licensed, there is no per-token fee layered on top; you pay only for the compute.


The verdict

GLM 5.2 is the most capable open-weights coding model you can verify. On Scale’s standardized SEAL harness it posts the top open score on SWE-Bench Pro (~62%), and on static evals it trades within a point or two of GPT-5 and Claude Opus 4.6 — all under the MIT license, the most permissive terms the industry offers. The 68.5% headline that launched a thousand posts in June was real but scaffold-dependent, and the August benchmark-integrity debate was right to force the distinction: on a common harness, the closed frontier still leads, by a few points rather than the old 5–10.

The hardware story has improved since launch. The full 753B model is still server-class for real work, but community quantizations now run it on a 256 GB M3 Ultra at ~22 tok/s (1-bit GGUF) and a 512 GB M3 Ultra at ~15 tok/s (4-bit MLX), both community-reported — a genuine first for a model this size on a desktop. For everyone else, GLM-4.5-Air delivers the GLM line on a 64 GB Mac at an estimated ~30 tok/s, with GLM-4.7-Flash covering the 32 GB tier.

The bigger story survives the asterisk. The era where “open” meant “good enough, but behind” is ending: according to the LLMCheck index, the standardized gap between the best open model and the closed frontier on real engineering work is now the smallest ever recorded — and the model holding that line is one you can download, fine-tune, and ship under MIT. For the ranked field, see the leaderboard and the August 2026 State of Open-Source Local LLMs report.