What Happened to "Qwen 4"?

Nothing happened to it, because it was never announced. After the Qwen 3.5 and 3.6 releases, much of the community assumed the next major version would be called "Qwen 4", and rumor posts filled in the blanks: a "Qwen 4.1 32B-A3B", a "Qwen 4 Coder", even a "Qwen 4 Preview". The invented specs were plausible precisely because they mirrored a real model — the circulated "32B-class MoE, ~3B active parameters, Apache 2.0, 256K context" description is essentially Qwen3.6-35B-A3B with the version number filed off.

During a July 2026 data incident, LLMCheck ingested several of those unverified entries into its catalog. An August 2026 reconciliation against official Qwen channels (the team's blog, GitHub, and Hugging Face organization) confirmed that no "Qwen 4" of any kind has ever been published, and the entries were removed from the leaderboard and benchmarks.

The tell, in hindsight: Alibaba announces every Qwen release through the same official channels within hours, with model cards and downloadable weights. A "major release" that exists only in aggregator tables and forum posts — with no weights anyone can pull — is a naming rumor, not a model.

The Real Roadmap: 3.5 → 3.6 → 3.8

Here is the verified state of the Qwen line as of August 2026:

Model Status License Coding score Mac fit
Qwen3.6-27B (dense) Shipped Apache 2.0 77.2% SWE-bench Verified 32GB+ (24GB tight)
Qwen3.6-35B-A3B (MoE) Shipped Apache 2.0 73.4% SWE-bench 24GB+
Qwen3.8-Max Hosted GA (August 2026) Proprietary, API-only — No
Qwen3.8-2.4T-A95B Open weights (August 2026) Custom (revenue threshold) 67.7 SWE-Bench Pro No (~397GB smallest quant)
Qwen3.8-27B Teased (August 2026) Unpublished Unpublished TBD
"Qwen 4 / 4.1 / 4 Coder" Never released — — —

Two structural notes. First, the version jump from 3.6 to 3.8 is Alibaba's own numbering — there was no intermediate 3.7 open release, which is part of why outsiders kept guessing "4 is next". Second, the 3.8 generation splits into a hosted flagship (Qwen3.8-Max) and open weights (the 2.4T-A95B), a pattern the earlier all-open Qwen generations did not have.

Qwen3.6: What to Run on a Mac

For local use, the Qwen3.6 family is the story, and it gives you a clean two-way choice.

Qwen3.6-27B — the capability pick

Qwen3.6-27B is a dense 27B model and the strongest open coder-generalist that fits on a normal Mac: 77.2% on SWE-bench Verified. According to the LLMCheck index it holds a Score of 72 — the current Mac-runnable #1 — at an estimated ~40 tok/s on an M5 Max. Being dense, every token pays the full 27B cost, so it wants memory bandwidth: it is happiest on Pro/Max-class chips with 32GB or more, though a 24GB machine can run a Q4 quant with care. Per-chip estimates: M5 Max, M4 Max, M4 Pro, M3 Max.

Qwen3.6-35B-A3B — the speed pick

Qwen3.6-35B-A3B is the mixture-of-experts sibling: roughly 35B total parameters with only ~3B active per token, under Apache 2.0. It gives up some accuracy (73.4% SWE-bench) but generates dramatically faster per token than the dense 27B, because the per-token compute tracks the 3B active set. If your Mac has 24GB, or your workload is long agentic loops where token throughput compounds, this is the one. Per-chip estimates: M5 Max, M4 Max, M4 Pro.

Run Qwen3.6-27B if…

You want the best answer quality a Mac can produce right now — coding accuracy above all. 77.2% SWE-bench Verified is the highest of any Mac-runnable open model in the LLMCheck index, and 32GB+ machines run it comfortably at Q4.

Run Qwen3.6-35B-A3B if…

You have 24GB of unified memory, or you run agents that burn thousands of tokens per task. The ~3B active-parameter MoE design trades a few benchmark points for much higher generation speed — and it keeps the same Apache 2.0 license.

Qwen3.8: Max, the 2.4T Open Weights, and the Teased 27B

The 3.8 generation arrived in August 2026 in three parts, and none of them changes the Mac recommendation yet.

Recommendations by Mac Tier

Installing the Real Qwens

Both Qwen3.6 models are available through the usual runners. Exact tags vary by registry — search "Qwen3.6" in Ollama or LM Studio and pick a Q4_K_M (or MLX 4-bit) build:

# Capability pick — the current Mac #1
ollama run qwen3.6:27b

# Speed pick for 24GB Macs
ollama run qwen3.6:35b-a3b

In LM Studio, search "Qwen3.6" and choose the quant that fits your memory tier. MLX-community conversions exist for both models and typically squeeze extra tok/s out of the unified-memory path on M-series chips.