What Happened to "DeepSeek R3"

During a July 2026 data incident, the LLMCheck catalog listed a "DeepSeek R3" as the reasoning-focused rival to GLM 5.2. When the catalog was reconciled against official sources in August 2026, no such release could be verified. DeepSeek's real lineage runs R1 → V3.2 / V3.2-Speciale → V4, and the V4 generation has two members: V4-Flash (the V4-Flash-0731 checkpoint, released in late July 2026, MIT, 284B-A13B) and V4 Pro (1.6T-A49B, GA in August 2026). Those are the models this page now compares against GLM 5.2 — and the comparison is more interesting than the fictional one, because one of these actually runs on a Mac.

Quick Verdict

If you only read one section, read this one. The two camps lead on different axes, and the tiebreaker for most readers is hardware, not benchmarks.

Choose GLM 5.2 for…

Peak open-weights agentic coding under a clean MIT license. On Scale's standardized SEAL harness it holds the top open SWE-Bench Pro score (~62%). The catch is scale: 753B parameters means the community has only run it on 256 GB+ M3 Ultra configurations, at extreme quantization.

Choose DeepSeek V4-Flash for…

Frontier-tier capability you can actually load on a 128 GB Mac. It scores 50 on the AA Intelligence Index — the top open tier — and self-reports 82.7 on Terminal Bench 2.1. Its 2-bit MLX build is 96.5 GB, with community-reported ~39 tok/s on an M5 Max.

According to the LLMCheck index, GLM 5.2 and DeepSeek V4-Flash are the two highest-capability open-weights entries in the catalog as of August 2026 — and V4-Flash is the only one of the pair that fits a 128 GB Mac.

The Lineup: One Flagship vs a Family

This is not a symmetric matchup. Zhipu fields one giant flagship; DeepSeek fields a family with two very different members. All three are sparse Mixture-of-Experts designs, which is what keeps their inference cost proportional to the small active parameter count rather than the headline total.

Spec GLM 5.2 DeepSeek V4-Flash DeepSeek V4 Pro
Developer Zhipu AI DeepSeek DeepSeek
Released June 2026 July 2026 August 2026 (GA)
Parameters 753B MoE 284B-A13B MoE 1.6T-A49B MoE
Weights license MIT MIT GA weights unpublished*
Mac floor 256 GB (1-bit, community) 128 GB (2-bit MLX) Not runnable

*V4 Pro's April preview weights are public and remain MIT-licensed; the GA build's weights were not published at launch, which makes the GA model an API proposition for now. The 13B active-parameter count on V4-Flash is the number to notice: it is what makes a 284B model move at interactive speed on Apple Silicon.

Scores, With Harness Notes

Every number below carries its provenance, because the single most important thing to understand about this matchup is that vendor scaffolds and standardized harnesses produce different numbers for the same model.

Model Score Benchmark Harness / provenance
GLM 5.2 68.5% SWE-Bench Pro Zhipu's own agent scaffold
GLM 5.2 ~62% SWE-Bench Pro Scale SEAL, standardized — open-weights lead
V4-Flash 50 AA Intelligence Index Artificial Analysis — top open tier
V4-Flash 82.7 Terminal Bench 2.1 Self-reported by DeepSeek
V4 Pro — — Independent harness results pending at GA

The GLM 5.2 rows are the lesson in miniature. Zhipu's famous 68.5% SWE-Bench Pro figure was produced with Zhipu's own scaffold; when Scale's SEAL team reran the top open models on a standardized harness, the best open score landed around 62% — with GLM 5.2 still leading the open field. Both statements are true: GLM 5.2 is the strongest open agentic coder on a neutral harness, and the headline number most people quote overstates it by several points.

V4-Flash's numbers need the same care. The 50 on the Artificial Analysis Intelligence Index is third-party and puts it in the top open tier; the 82.7 on Terminal Bench 2.1 is DeepSeek's own report and awaits independent replication. V4 Pro went GA in August 2026 with server-class positioning, and the LLMCheck index will list harness-verified numbers as they appear rather than reprinting launch claims.

License & Weights: Where MIT Ends

GLM 5.2 and V4-Flash are both genuinely open: MIT-licensed weights, free to download, fine-tune, self-host, and ship commercially. On the LLMCheck methodology, MIT earns the full 10/10 license score — no user caps, no acceptable-use bolt-ons, no revenue gates.

V4 Pro is the asterisk. The April preview weights are public under MIT and stay that way, but the improved GA build that launched in August 2026 shipped without published weights. If GA weights arrive later, V4 Pro becomes the largest MIT model in the catalog; until then, treat it as an API model with an open ancestor. For anyone whose requirement is "weights on my own hardware," today's real choice is GLM 5.2 versus V4-Flash.

Mac Reality: 128 GB vs 256 GB

This is where the matchup stops being academic. Neither flagship was designed for a Mac — but the community has forced both onto Apple Silicon, at very different price points.

DeepSeek V4-Flash — the 128 GB Mac story of the summer

The 2-bit MLX quantization of V4-Flash weighs 96.5 GB, which fits a 128 GB Mac with room for context. Community-reported throughput is ~39 tok/s on an M5 Max and ~27 tok/s on an M3 Max, the latter via antirez's ds4 Metal engine. Tooling is moving fast: llama.cpp merged DSpark speculative-decoding support in early August 2026, though V4-Flash is not in Ollama yet — MLX is the path today. Chip-level detail lives at DeepSeek V4-Flash on M5 Max and on M3 Max.

2-bit is an aggressive quantization. Community reports so far are positive for coding and agentic use, but expect some fidelity loss versus the full-precision model on the hardest reasoning — and treat all current tok/s figures as community numbers, not vendor specs.

GLM 5.2 — possible, at the very top of the Mac range

At 753B total parameters, GLM 5.2 only fits the largest Apple Silicon configurations, and only at extreme quantization: community-reported runs put a 1-bit GGUF at ~22 tok/s on a 256 GB M3 Ultra and a 4-bit MLX build at ~15 tok/s on a 512 GB M3 Ultra (see GLM 5.2 on M3 Ultra). Below that tier, the practical GLMs are GLM-4.5-Air on 64 GB Macs and GLM-4.7-Flash on 24–32 GB machines — our GLM-4.5-Air guide covers both.

⚡ Run it at full size — rent a GPU

The full GLM 5.2 and DeepSeek V4 Pro are server-class. To run them unquantized, rent a datacenter GPU by the minute on Vast.ai — often 5–6× cheaper than AWS or GCP, with H100s and B200s available on demand.

Vast.ai referral link — we may earn a small commission at no extra cost to you. It never influences our reviews or rankings.

The Verdict

The real showdown has a cleaner answer than the fictional one did, because hardware breaks the tie.

The deeper story of August 2026 is that "open frontier" now splits into two questions: which model is strongest, and which strongest model fits your machine. GLM 5.2 answers the first on a neutral harness; V4-Flash is, according to the LLMCheck index, the best current answer to the second — see how both rank against everything else on the leaderboard.