The 30-Day Recap (TL;DR)
- The M6 is measured. A community benchmark of a 32 GB Mac mini M6, with its raw data published, put Qwen 3.6-35B-A3B at 64 tok/s and dense Gemma 4 31B at 8.
- Small MoEs move to the top. Six fully specified runs showed small mixture-of-experts models decoding two to four times faster than the index estimated. After the refit, KAT-Coder-V2.5 (Score 82) and Qwen 3.6-35B-A3B (81) lead; Qwen3.8-27B is #8 and the highest-ranked dense model.
- Two corrections. The index carried the M5 Pro and M5 Max at 273 and 600 GB/s on some configurations (Apple: 307 and 614), and it listed a “Mac Studio M4 Ultra” that Apple never made. Both are fixed, with dated notes on every affected page.
- Two models added: OpenAI’s GPT-oss 20B and Google’s Gemma 4 12B, both measured on the M6.
- Runtimes: stock llama.cpp runs GLM-5.3-Flash since 30 September; Ollama 0.40 (4 October) runs supported models on MLX by default; DeepSeek V4.1 Flash still has no upstream runtime.
- Announced, not shipped: Qwen 4, including an open-weight 27B. Muse Spark 1.2 weights remain a promise.
The First Mac mini M6 Measurements
Apple’s new entry Mac shipped on 22 September, and the first methodical numbers came from a community benchmark of a 32 GB Mac mini M6: LM Studio, the median of three runs, an English-prose prompt, and the raw JSON published alongside the script. That last part is why the index can use it — every figure is checkable.
| Model | Type | Index estimate before | Measured (M6, 32 GB) |
|---|---|---|---|
| Qwen 3.6-35B-A3B | MoE, 3B active | 14 tok/s | 64 tok/s |
| Gemma 4 26B-A4B | MoE, 4B active | 14 | 48 |
| Nemotron 3.5 Lightning | MoE, 3B active | 17 | 45 |
| GPT-oss 20B | MoE, 3.6B active | — (not yet listed) | 45 |
| Gemma 4 12B | Dense | 19 (formula) | 19 |
| Muse Glimmer 30B | Dense | 8 | 9 |
| Gemma 4 31B | Dense | 8 (formula) | 8 |
The dense rows are the reassuring half: the bandwidth formula landed within a token per second on all three. The mixture-of-experts rows are the other half — off by a factor of 2.6 to 4.6. An M5 Max community run of Qwen 3.6-35B-A3B told the same story: 126 tok/s against an estimate of about 50. For buyers the practical point is simple: on a $1,399 Mac mini, a 35B mixture-of-experts model generates faster than most people read.
Why the Leaderboard Reshuffled
A mixture-of-experts model stores all its weights but reads only the active ones per token, so its speed should track active-parameter bytes. The index had scaled that by a flat 0.42, fitted to one 2-bit DeepSeek V4 Flash run. Six fully specified plain-decode measurements of small MoEs — the M6 runs above, the M5 Max run, and two M4 Max runs from a published comparison (arXiv 2601.19139) — fit the dense formula with a larger fixed cost per token, about 5 ms instead of 1.5, with a mean error near 12%.
What changed: small MoEs (40B total or less) now use that fitted model; large MoEs keep the conservative 0.42 factor, because the only fully specified large-MoE runs still support it; and a model with its own measurement has every other chip scaled from it. The methodology page has the formulas and all six data points.
Because speed points cap at 25 — reached at 100 tok/s on the reference chip — small MoEs that now estimate above 100 tok/s on an M5 Max collected the full 25. That is what moved KAT-Coder-V2.5, Qwen 3.6-35B-A3B, Nemotron 3.5 Lightning and Gemma 4 26B-A4B past the dense 27B models. Capability scores did not change; Qwen3.8-27B still has one of the highest capability scores a 24 GB Mac can run, at about 29 tok/s on an M5 Max.
Two Corrections
Apple’s M5 bandwidths. Apple specifies 307 GB/s for the M5 Pro and 614 GB/s for the M5 Max (460 on the 32-core-GPU version). The index’s MacBook Pro configurations still carried 273 and 600 from before the chips launched, and because of how the estimator read its chip table, those values won. M5 Pro estimates rise by up to 12%; M5 Max estimates are unchanged, because the estimator’s efficiency constant was refit with the correct bandwidth; other chips move about 2%. Binned configurations — the 36 GB Mac Studio M5 Max at 460 GB/s and the 16 GB Mac mini M6 at 153 — now use their own figures instead of the full chip’s.
The Mac Studio M4 Ultra never existed. Apple confirmed in March 2025 that the M4 Max has no UltraFusion connector, and paired it with the M3 Ultra in that year’s Mac Studio. The index nevertheless listed a “Mac Studio M4 Ultra” — with 18 estimated speed rows, two Best-by-Mac pages and Amazon links. It has been removed; its pages redirect to the M5 Ultra, and the articles that cited it carry dated corrections. A validation rule now rejects any chip Apple did not make.
The same review fixed the Mac Advisor, which had scaled every model as if the index’s reference speeds were M3 Max figures — overstating an M5 Max by about 1.5× and an Ultra by 2× — and the leaderboard’s speed column, which from 23 September ranked MiMo-V2.6-Flash #1 on an M5 Ultra estimate shown as an M5 Max figure. Models too big for any M5 Max are now ranked on the M5 Ultra and tagged as such.
New in the Index: GPT-oss 20B and Gemma 4 12B
GPT-oss 20B (OpenAI, Apache 2.0) is a 21B mixture-of-experts model with 3.6B active parameters, shipped in MXFP4 so it fits a 16 GB Mac. Its model card reports 53.2 on SWE-bench Verified; the M6 run measured 45 tok/s. It had been missing from the catalog while its 120B sibling was listed — and that sibling, GPT-oss 120B, was being costed as a dense model; it is a 117B MoE with 5.1B active, and its figures are corrected.
Gemma 4 12B (Google, Apache 2.0) is the dense member of the Gemma 4 family: text, image and audio input, a 256K context, and model-card scores close to the 26B-A4B (MMLU-Pro 77.2, GPQA Diamond 78.8) at about 7 GB at 4-bit. It needs a 16 GB Mac; the M6 measured 19 tok/s.
Runtimes: llama.cpp, Ollama, DeepSeek V4.1
- GLM-5.3-Flash in stock llama.cpp. The support pull request merged on 30 September (release b11279). Use GGUFs converted for the merged
glm5-nextarchitecture — files made for the earlier support branch will not load — and expect multi-token prediction in a separate pull request. It still needs a 128 GB Mac at 3-bit. Setup guide. - Ollama goes MLX-first. Version 0.40 (4 October) runs every architecture its MLX engine supports on MLX by default on Apple Silicon; 0.35 added an API for decision models (below).
- DeepSeek V4.1 Flash: still no upstream runtime. Its causal encoder-decoder design is not in llama.cpp, mlx-lm or transformers yet. Community ports exist — custom MLX loaders, a llama.cpp branch — and an upstream conversion pull request is open, so the index keeps it without a Mac speed until a mainstream runtime lands.
Decision Models Arrive
A new class of small model appeared on 1 October: decision models, which return a choice, a probability or a score instead of text — for routing, triage, tool selection and guardrails. Cloudflare open-sourced Clef and Clef-flash under Apache 2.0, and AWS’s Strands team released Strands Decider 2B with its training data. Ollama serves the format through a dedicated endpoint. They are not chat models, so they do not enter the leaderboard, but at 2B they run comfortably on any Mac — a useful companion to the agent models above.
Open-Source Top 10 (October 2026)
| # | Model | Score | License | Min RAM | Speed (M5 Max unless noted) |
|---|---|---|---|---|---|
| 1 | KAT-Coder-V2.5 | 82 | Apache 2.0 | 24 GB | 117 tok/s est. |
| 2 | Qwen 3.6-35B-A3B | 81 | Apache 2.0 | 24 GB | 126 tok/s community |
| 3 | Nemotron 3.5 Lightning | 79 | OpenMDW | 18 GB | 103 tok/s est. |
| 4 | Gemma 4 26B-A4B | 78 | Apache 2.0 | 18 GB | 106 tok/s est. |
| 5 | Laguna XS 2.1 | 76 | OpenMDW | 24 GB | 117 tok/s est. |
| 6 | MiMo-V2.6-Flash | 72 | MIT | 170 GB | 42 tok/s est. (M5 Ultra) |
| 7 | Nemotron-Cascade 2 | 71 | NVIDIA Open | 18 GB | 117 tok/s est. |
| 8 | Qwen3.8-27B | 71 | Apache 2.0 | 19 GB | 29 tok/s est. |
| 9 | DeepSeek V4 Flash | 70 | MIT | 120 GB | 39 tok/s community |
| 10 | DeepSeek V4 Flash Vision-Exp | 70 | MIT | 120 GB | 39 tok/s est. |
Every speed figure carries its provenance on the leaderboard. KAT-Coder-V2.5’s #1 rests on an estimate from the refit model; Qwen 3.6-35B-A3B’s #2 on a community measurement. A measured KAT-Coder figure is the obvious next test of the refit.
By Mac Tier — What to Run Today (October 2026)
| Memory | Mac | Run | Speed |
|---|---|---|---|
| 16 GB | Mac mini M6, MacBook Air | GPT-oss 20B; Gemma 4 12B | 45 and 19 tok/s measured on an M6 |
| 24–32 GB | Mac mini M6, MacBook Air M4 | Qwen 3.6-35B-A3B; Gemma 4 26B-A4B; KAT-Coder-V2.5 | 64 and 48 measured on an M6; KAT ~56 est. |
| 48–64 GB | Mac mini M5 Pro, MacBook Pro M5 Pro | Qwen3-Coder-Next; 70B dense at 4-bit | ~49 and ~6 tok/s est. |
| 128 GB | Mac Studio / MacBook Pro M5 Max | DeepSeek V4 Flash (2-bit); Qwen3.8-Flash-Next | 39 community; ~49 est. |
| 256–512 GB | Mac Studio M5 Ultra / M3 Ultra | MiMo-V2.6-Flash; MiniMax M2.5; GLM-5.3 at 512 GB | ~42, ~57 and ~17 est. on an M5 Ultra |
Exact figures for 42 Mac configurations are on the Best-by-Mac pages; the Mac Advisor now reads the same numbers.
What’s Coming Next Month
Speculative section. Each item carries its status; nothing here enters the index until weights exist on a primary source.
- Announced Qwen 4. Alibaba named four tiers at Apsara on 22 September — Max, Plus, Flash and an open-weight 27B — and said the family is still in training. No weights, license or date. Qwen3.8-Flash-Next remains the only public preview of the architecture.
- Announced Muse Spark 1.2 weights. Promised on 10 August and again on 2 September; still not published as of 5 October.
- Announced The 512 GB Mac Studio M5 Ultra ships from late October — the first Mac that can hold GLM-5.3 at 4-bit at 1.2 TB/s.
- Watch Measurements the index still lacks: a published M6 figure from a press review, a plain-decode (no MTP) figure for Qwen3.8-Flash-Next, a measured KAT-Coder-V2.5, and GLM-5.3-Flash’s multi-token prediction in llama.cpp.
Frequently Asked Questions
What is the best open-source local LLM for a Mac in October 2026?
According to the LLMCheck index, KAT-Coder-V2.5 (Score 82) and Qwen 3.6-35B-A3B (Score 81) — both 30B-class mixture-of-experts models that fit a 24 GB Mac. Qwen 3.6-35B-A3B has a community-measured 126 tok/s on an M5 Max. The strongest dense model is Qwen3.8-27B, at #8.
How fast is the Mac mini M6 for local LLMs?
A community benchmark of a 32 GB Mac mini M6 measured Qwen 3.6-35B-A3B at 64 tok/s, Gemma 4 26B-A4B at 48, GPT-oss 20B at 45 and dense Gemma 4 31B at 8 (LM Studio, median of three runs). Small mixture-of-experts models are the M6's sweet spot.
Why did the LLMCheck leaderboard change in October 2026?
The index refit its mixture-of-experts speed estimates to six measured runs on three chips, which showed small MoEs running two to four times faster than the old flat factor predicted. Several now estimate above 100 tok/s on an M5 Max, where speed points cap, which moved them past dense models.
Does GLM-5.3-Flash run in llama.cpp now?
Yes. Upstream llama.cpp merged support on 30 September 2026 (release b11279). Use GGUFs converted for the merged glm5-next architecture; files from the earlier support branch will not load. It needs a 128 GB Mac at 3-bit.
Is Qwen 4 out?
No. Alibaba announced the Qwen 4 family at Apsara on 22 September — Max, Plus, Flash and an open-weight 27B — and said it is still in training. No weights, license or release date have been published; Qwen3.8-Flash-Next is the closest preview of its architecture.
What is the M5 Max's memory bandwidth?
614 GB/s on the 40-core GPU and 460 GB/s on the 32-core version; the M5 Pro has 307 GB/s. Until 5 October the LLMCheck index used 600 and 273 for its MacBook Pro configurations; estimates for M5 Pro Macs rose by up to 12% with the correction.