What Happened to "Qwen 4"?
Nothing happened to it, because it was never announced. After the Qwen 3.5 and 3.6 releases, much of the community assumed the next major version would be called "Qwen 4", and rumor posts filled in the blanks: a "Qwen 4.1 32B-A3B", a "Qwen 4 Coder", even a "Qwen 4 Preview". The invented specs were plausible precisely because they mirrored a real model — the circulated "32B-class MoE, ~3B active parameters, Apache 2.0, 256K context" description is essentially Qwen3.6-35B-A3B with the version number filed off.
During a July 2026 data incident, LLMCheck ingested several of those unverified entries into its catalog. An August 2026 reconciliation against official Qwen channels (the team's blog, GitHub, and Hugging Face organization) confirmed that no "Qwen 4" of any kind has ever been published, and the entries were removed from the leaderboard and benchmarks.
The tell, in hindsight: Alibaba announces every Qwen release through the same official channels within hours, with model cards and downloadable weights. A "major release" that exists only in aggregator tables and forum posts — with no weights anyone can pull — is a naming rumor, not a model.
The Real Roadmap: 3.5 → 3.6 → 3.8
Here is the verified state of the Qwen line as of August 2026:
| Model | Status | License | Coding score | Mac fit |
|---|---|---|---|---|
| Qwen3.6-27B (dense) | Shipped | Apache 2.0 | 77.2% SWE-bench Verified | 32GB+ (24GB tight) |
| Qwen3.6-35B-A3B (MoE) | Shipped | Apache 2.0 | 73.4% SWE-bench | 24GB+ |
| Qwen3.8-Max | Hosted GA (August 2026) | Proprietary, API-only | — | No |
| Qwen3.8-2.4T-A95B | Open weights (August 2026) | Custom (revenue threshold) | 67.7 SWE-Bench Pro | No (~397GB smallest quant) |
| Qwen3.8-27B | Teased (August 2026) | Unpublished | Unpublished | TBD |
| "Qwen 4 / 4.1 / 4 Coder" | Never released | — | — | — |
Two structural notes. First, the version jump from 3.6 to 3.8 is Alibaba's own numbering — there was no intermediate 3.7 open release, which is part of why outsiders kept guessing "4 is next". Second, the 3.8 generation splits into a hosted flagship (Qwen3.8-Max) and open weights (the 2.4T-A95B), a pattern the earlier all-open Qwen generations did not have.
Qwen3.6: What to Run on a Mac
For local use, the Qwen3.6 family is the story, and it gives you a clean two-way choice.
Qwen3.6-27B — the capability pick
Qwen3.6-27B is a dense 27B model and the strongest open coder-generalist that fits on a normal Mac: 77.2% on SWE-bench Verified. According to the LLMCheck index it holds a Score of 72 — the current Mac-runnable #1 — at an estimated ~40 tok/s on an M5 Max. Being dense, every token pays the full 27B cost, so it wants memory bandwidth: it is happiest on Pro/Max-class chips with 32GB or more, though a 24GB machine can run a Q4 quant with care. Per-chip estimates: M5 Max, M4 Max, M4 Pro, M3 Max.
Qwen3.6-35B-A3B — the speed pick
Qwen3.6-35B-A3B is the mixture-of-experts sibling: roughly 35B total parameters with only ~3B active per token, under Apache 2.0. It gives up some accuracy (73.4% SWE-bench) but generates dramatically faster per token than the dense 27B, because the per-token compute tracks the 3B active set. If your Mac has 24GB, or your workload is long agentic loops where token throughput compounds, this is the one. Per-chip estimates: M5 Max, M4 Max, M4 Pro.
Run Qwen3.6-27B if…
You want the best answer quality a Mac can produce right now — coding accuracy above all. 77.2% SWE-bench Verified is the highest of any Mac-runnable open model in the LLMCheck index, and 32GB+ machines run it comfortably at Q4.
Run Qwen3.6-35B-A3B if…
You have 24GB of unified memory, or you run agents that burn thousands of tokens per task. The ~3B active-parameter MoE design trades a few benchmark points for much higher generation speed — and it keeps the same Apache 2.0 license.
Qwen3.8: Max, the 2.4T Open Weights, and the Teased 27B
The 3.8 generation arrived in August 2026 in three parts, and none of them changes the Mac recommendation yet.
- Qwen3.8-Max went GA as a hosted model in early August. It is API-only — no weights — so it sits outside the scope of the LLMCheck index.
- Qwen3.8-2.4T-A95B open weights followed in August: 2.4 trillion total parameters with 95B active, scoring 67.7 on SWE-Bench Pro and 92.6 on GPQA. Two catches. The smallest published quantization is roughly 397GB, so no Mac can hold it. And it ships under a custom license with a revenue threshold — the first Qwen open release to move away from Apache 2.0.
- Qwen3.8-27B has been teased for later in August 2026. Its license and benchmarks are unpublished, so the LLMCheck index does not list it. If it lands under a permissive license, it becomes the obvious challenger to Qwen3.6-27B — until then, treat every "Qwen3.8-27B benchmark" you see with the same skepticism this page exists to teach.
Recommendations by Mac Tier
- 8–16GB — Skip the 27B-class entirely. Qwen3.5-9B remains the sensible Qwen at this tier (per-chip estimates), or see the Best by Mac pages for your exact machine.
- 24GB — Qwen3.6-35B-A3B at Q4. The MoE design is what makes a 35B-class model livable in this footprint.
- 32–64GB — Qwen3.6-27B. This is the sweet spot for the current Mac #1, with room for context and your editor.
- 64GB+ — Qwen3.6-27B for general work; for coding specifically, the 80B-A3B Qwen3-Coder-Next also becomes an option at this tier.
Installing the Real Qwens
Both Qwen3.6 models are available through the usual runners. Exact tags vary by registry — search "Qwen3.6" in Ollama or LM Studio and pick a Q4_K_M (or MLX 4-bit) build:
ollama run qwen3.6:27b
# Speed pick for 24GB Macs
ollama run qwen3.6:35b-a3b
In LM Studio, search "Qwen3.6" and choose the quant that fits your memory tier. MLX-community conversions exist for both models and typically squeeze extra tok/s out of the unified-memory path on M-series chips.