What Happened to "Qwen 4.1" and "Llama 5"

During a July 2026 data incident, this page compared GLM 5.2 against a "Qwen 4.1 32B-A3B" and a "Llama 5 405B". Neither could be verified against any official source when the catalog was reconciled in August 2026. The real Qwen lineage runs 3.5 → 3.6 → 3.8, and Meta's open flagship remains Llama 4 Maverick — the next Llama generation is not expected before 2027, and Meta's actual open release of August 2026 is Muse Glimmer 30B (Apache 2.0), a 30B dense multimodal agent model.

The corrected three-way is the real state of the open frontier: GLM 5.2 (Zhipu AI, June 2026), Qwen3.8-Max — Alibaba's hosted flagship, whose open-weights sibling Qwen3.8-2.4T-A95B shipped in August 2026 — and Kimi K3 (Moonshot AI, weights released in late July 2026).

The Three-Way Verdict at a Glance

These three flagships are not competing for the same crown. One wins on license, one on published scores, one on sheer scale — and none of them wins on your Mac.

License King

GLM 5.2

Zhipu's 753B MoE is the only one of the three under a clean MIT license, and it holds the top open-weights SWE-Bench Pro score on Scale's standardized SEAL harness (~62%). The only flagship here the community has run on a Mac at all — 256 GB M3 Ultra territory.

Score King

Qwen3.8-Max

The open-weights 2.4T-A95B posts the highest published numbers of the three: 67.7 SWE-Bench Pro and 92.6 GPQA, both vendor-reported. The catch: a custom revenue-threshold license — the first Qwen flagship off Apache 2.0 — and a smallest quant of ~397 GB.

Scale King

Kimi K3

Moonshot's 2.8T MoE is the largest open-weight model ever released — ~1.56 TB of weights. It is the #1 open model on LiveBench Coding and #3 on the AA Intelligence Index, under an MIT-like license with a $20M model-as-a-service revenue gate. Not Mac-runnable, period.

Scores, With Harness Notes

Every headline number below comes from a different measurement setup, so each row carries its provenance. That is not pedantry — the gap between a vendor scaffold and a standardized harness can be larger than the gap between the models.

Model Score Benchmark Harness / provenance
GLM 5.2 68.5 SWE-Bench Pro Zhipu's own agent scaffold
GLM 5.2 ~62 SWE-Bench Pro Scale SEAL, standardized — open-weights lead
Qwen3.8-2.4T 67.7 SWE-Bench Pro Vendor-reported (Alibaba's setup)
Qwen3.8-2.4T 92.6 GPQA Vendor-reported
Kimi K3 #1 open LiveBench Coding Public leaderboard
Kimi K3 #3 open AA Intelligence Index Artificial Analysis

Read the two GLM 5.2 rows together and the lesson writes itself: Zhipu's famous 68.5 was produced on Zhipu's own scaffold, while Scale's standardized SEAL rerun puts the best open score near 62 — with GLM 5.2 still on top of the open field. Qwen's 67.7 and Zhipu's 68.5 are therefore not directly comparable: both are vendor numbers from different setups, and Qwen3.8's standardized SEAL results were not yet published when this page was updated. Kimi K3's rankings come from third-party boards, which makes them the most conservative claims in the table.

According to the LLMCheck index, the honest summary is: Qwen3.8-2.4T has the highest published scores, GLM 5.2 has the strongest independently standardized result, and Kimi K3 has the strongest third-party coding ranking. Anyone declaring a single "best open model" without naming the harness is selling something.

License Three-Way: MIT vs Revenue Gates

This comparison used to be the boring section. In August 2026 it is the decisive one, because two of the three flagships shipped with custom terms.

License Trait GLM 5.2 (MIT) Qwen3.8-2.4T (custom) Kimi K3 (custom)
Commercial use Yes, unrestricted Yes, below threshold Yes, below gate
Revenue / usage gate None Revenue threshold $20M/yr MaaS revenue
Weights downloadable Yes Yes (~397 GB min quant) Yes (~1.56 TB)
LLMCheck license tier Top (10/10) Reduced (custom terms) Reduced (custom terms)

If your requirement is zero license ambiguity, GLM 5.2 wins this section outright. The Qwen move matters beyond this comparison: Apache 2.0 was the Qwen line's calling card, and the 3.8 flagship abandoning it is the clearest sign yet that frontier-scale open weights are drifting toward conditional licenses.

Mac Reality: None of the Above

Here is the section that separates this page from a press-release roundup: none of these three flagships is a practical Mac model.

What Mac users should actually run

The frontier trickles down fast, and the local winners in August 2026 mostly come from these same labs:

For a ranked list matched to your exact machine, the Best LLM by Mac hub covers every chip and RAM tier.

⚡ Run it at full size — rent a GPU

All three flagships are server-class. To run any of them unquantized, rent a datacenter GPU by the minute on Vast.ai — often 5–6× cheaper than AWS or GCP, with H100s and B200s available on demand.

Vast.ai referral link — we may earn a small commission at no extra cost to you. It never influences our reviews or rankings.

Use-Case Picks

Matching the flagship to the job removes most of the ambiguity:

The Verdict

The open frontier in August 2026 splits three ways, and the split is about terms and tonnage as much as capability:

If you take one thing away: the open frontier's top end has left consumer hardware behind, and the licenses are following. The models Mac users actually run — Qwen3.6-27B, DeepSeek V4-Flash, Muse Glimmer 30B — are distillations and mid-sizers from the same labs, and according to the LLMCheck index that is where the local value lives. The leaderboard ranks all of them side by side.