What Happened to "Qwen 4.1" and "Llama 5"
During a July 2026 data incident, this page compared GLM 5.2 against a "Qwen 4.1 32B-A3B" and a "Llama 5 405B". Neither could be verified against any official source when the catalog was reconciled in August 2026. The real Qwen lineage runs 3.5 → 3.6 → 3.8, and Meta's open flagship remains Llama 4 Maverick — the next Llama generation is not expected before 2027, and Meta's actual open release of August 2026 is Muse Glimmer 30B (Apache 2.0), a 30B dense multimodal agent model.
The corrected three-way is the real state of the open frontier: GLM 5.2 (Zhipu AI, June 2026), Qwen3.8-Max — Alibaba's hosted flagship, whose open-weights sibling Qwen3.8-2.4T-A95B shipped in August 2026 — and Kimi K3 (Moonshot AI, weights released in late July 2026).
The Three-Way Verdict at a Glance
These three flagships are not competing for the same crown. One wins on license, one on published scores, one on sheer scale — and none of them wins on your Mac.
GLM 5.2
Zhipu's 753B MoE is the only one of the three under a clean MIT license, and it holds the top open-weights SWE-Bench Pro score on Scale's standardized SEAL harness (~62%). The only flagship here the community has run on a Mac at all — 256 GB M3 Ultra territory.
Qwen3.8-Max
The open-weights 2.4T-A95B posts the highest published numbers of the three: 67.7 SWE-Bench Pro and 92.6 GPQA, both vendor-reported. The catch: a custom revenue-threshold license — the first Qwen flagship off Apache 2.0 — and a smallest quant of ~397 GB.
Kimi K3
Moonshot's 2.8T MoE is the largest open-weight model ever released — ~1.56 TB of weights. It is the #1 open model on LiveBench Coding and #3 on the AA Intelligence Index, under an MIT-like license with a $20M model-as-a-service revenue gate. Not Mac-runnable, period.
Scores, With Harness Notes
Every headline number below comes from a different measurement setup, so each row carries its provenance. That is not pedantry — the gap between a vendor scaffold and a standardized harness can be larger than the gap between the models.
| Model | Score | Benchmark | Harness / provenance |
|---|---|---|---|
| GLM 5.2 | 68.5 | SWE-Bench Pro | Zhipu's own agent scaffold |
| GLM 5.2 | ~62 | SWE-Bench Pro | Scale SEAL, standardized — open-weights lead |
| Qwen3.8-2.4T | 67.7 | SWE-Bench Pro | Vendor-reported (Alibaba's setup) |
| Qwen3.8-2.4T | 92.6 | GPQA | Vendor-reported |
| Kimi K3 | #1 open | LiveBench Coding | Public leaderboard |
| Kimi K3 | #3 open | AA Intelligence Index | Artificial Analysis |
Read the two GLM 5.2 rows together and the lesson writes itself: Zhipu's famous 68.5 was produced on Zhipu's own scaffold, while Scale's standardized SEAL rerun puts the best open score near 62 — with GLM 5.2 still on top of the open field. Qwen's 67.7 and Zhipu's 68.5 are therefore not directly comparable: both are vendor numbers from different setups, and Qwen3.8's standardized SEAL results were not yet published when this page was updated. Kimi K3's rankings come from third-party boards, which makes them the most conservative claims in the table.
According to the LLMCheck index, the honest summary is: Qwen3.8-2.4T has the highest published scores, GLM 5.2 has the strongest independently standardized result, and Kimi K3 has the strongest third-party coding ranking. Anyone declaring a single "best open model" without naming the harness is selling something.
License Three-Way: MIT vs Revenue Gates
This comparison used to be the boring section. In August 2026 it is the decisive one, because two of the three flagships shipped with custom terms.
- GLM 5.2 — MIT. Genuinely unrestricted: commercial use, modification, redistribution, closed-product embedding. No caps, no gates. Top license score (10/10) on the LLMCheck methodology.
- Qwen3.8-2.4T — custom revenue-threshold license. The first Qwen flagship to move off Apache 2.0. Weights are downloadable and commercial use is permitted, but a revenue threshold in the terms triggers separate licensing for the largest deployers. Read the exact text before building a business on it.
- Kimi K3 — the "Kimi K3 License". MIT-like in day-to-day effect, with one custom clause: once model-as-a-service revenue built on K3 exceeds $20M per year, a separate agreement with Moonshot is required. For self-hosting and internal use it behaves like MIT; for API resellers it is a real gate.
| License Trait | GLM 5.2 (MIT) | Qwen3.8-2.4T (custom) | Kimi K3 (custom) |
|---|---|---|---|
| Commercial use | Yes, unrestricted | Yes, below threshold | Yes, below gate |
| Revenue / usage gate | None | Revenue threshold | $20M/yr MaaS revenue |
| Weights downloadable | Yes | Yes (~397 GB min quant) | Yes (~1.56 TB) |
| LLMCheck license tier | Top (10/10) | Reduced (custom terms) | Reduced (custom terms) |
If your requirement is zero license ambiguity, GLM 5.2 wins this section outright. The Qwen move matters beyond this comparison: Apache 2.0 was the Qwen line's calling card, and the 3.8 flagship abandoning it is the clearest sign yet that frontier-scale open weights are drifting toward conditional licenses.
Mac Reality: None of the Above
Here is the section that separates this page from a press-release roundup: none of these three flagships is a practical Mac model.
- Kimi K3 — ~1.56 TB of weights. No Mac, and no single consumer machine of any kind, can load it. It is a server-cluster model full stop.
- Qwen3.8-2.4T — smallest quantization ~397 GB. Even a 512 GB M3 Ultra would have no working headroom; the LLMCheck index does not list it as Mac-practical.
- GLM 5.2 — the only one with community Mac runs on record: a 1-bit GGUF at ~22 tok/s on a 256 GB M3 Ultra and a 4-bit MLX build at ~15 tok/s on a 512 GB M3 Ultra (both community-reported; see GLM 5.2 on M3 Ultra).
What Mac users should actually run
The frontier trickles down fast, and the local winners in August 2026 mostly come from these same labs:
- Qwen3.6-27B — the current LLMCheck Mac #1 (Score 72): a dense 27B with 77.2% SWE-bench Verified, at an estimated ~40 tok/s on an M5 Max. Its MoE sibling Qwen3.6-35B-A3B (Apache 2.0) is the faster everyday pick.
- DeepSeek V4-Flash — frontier-tier on a 128 GB Mac: its 2-bit MLX build is 96.5 GB, with community-reported ~39 tok/s on an M5 Max.
- Muse Glimmer 30B — Meta's real August 2026 release: Apache 2.0, ~17–18 GB at Q4 for 24–32 GB Macs, vendor-reported 26.6 tok/s on an M5 Max (50.2 with its bundled DFlash drafter).
- Inkling-Small (Thinking Machines Lab, July 2026, Apache 2.0, 276B-A12B) — holds the open-weight SWE-bench Verified record at 80.2% and fits a 128 GB Mac Studio via a 2-bit MLX build (88.4 GB); MLX-only for now.
For a ranked list matched to your exact machine, the Best LLM by Mac hub covers every chip and RAM tier.
All three flagships are server-class. To run any of them unquantized, rent a datacenter GPU by the minute on Vast.ai — often 5–6× cheaper than AWS or GCP, with H100s and B200s available on demand.
Vast.ai referral link — we may earn a small commission at no extra cost to you. It never influences our reviews or rankings.
Use-Case Picks
Matching the flagship to the job removes most of the ambiguity:
- Commercial deployment with legal scrutiny — GLM 5.2. MIT with no gates is the only friction-free option among the three.
- Peak published capability via API or your own cluster — Qwen3.8-Max (hosted) or the 2.4T open weights, if the revenue-threshold terms fit your business.
- Agentic coding at frontier scale — Kimi K3 if you have the infrastructure (LiveBench Coding #1 open); GLM 5.2 if you want the standardized-harness leader under MIT.
- Anything on a Mac — none of the three. Run Qwen3.6-27B, DeepSeek V4-Flash, or Muse Glimmer 30B instead, per the section above.
- Watching the horizon — Qwen3.8-27B, teased for later in August 2026 with license and benchmarks unpublished. If it stays Mac-sized, it could reset the local rankings.
The Verdict
The open frontier in August 2026 splits three ways, and the split is about terms and tonnage as much as capability:
- GLM 5.2 wins on license and standardized results. MIT, no gates, and the top open SWE-Bench Pro score on Scale's SEAL harness. It is also the only flagship here with any Mac existence at all — barely, at 256 GB and 1-bit.
- Qwen3.8-Max wins on published scores. 67.7 SWE-Bench Pro and 92.6 GPQA (vendor-reported) lead the table, but the custom revenue-threshold license and 397 GB minimum footprint make it an API-or-cluster proposition.
- Kimi K3 wins on scale and third-party coding rank. The largest open-weight model ever, #1 open on LiveBench Coding — and, at 1.56 TB, the clearest possible reminder that "open weights" and "runnable" are different words.
If you take one thing away: the open frontier's top end has left consumer hardware behind, and the licenses are following. The models Mac users actually run — Qwen3.6-27B, DeepSeek V4-Flash, Muse Glimmer 30B — are distillations and mid-sizers from the same labs, and according to the LLMCheck index that is where the local value lives. The leaderboard ranks all of them side by side.