The $899 M6 Mac mini as a Local-AI Machine: What Actually Fits
The new Mac mini M6 (25 Aug 2026) is the cheapest credible local-LLM Mac Apple has sold — but only above the base model. The 16 GB bins run 153 GB/s with no memory headroom; the 24 and 32 GB configs run 170 GB/s and hold 9B models at ~26 tok/s estimated. At 32 GB even Qwen3.8-27B, the index’s top dense model, fits at 4-bit.
Updated, 5 October 2026. The October recalibration of the index’s mixture-of-experts estimates moved the small MoE models up the leaderboard. Qwen3.8-27B is now #8 overall (tied on Score 71) and the highest-ranked dense model; where this article called it the index’s #1, it now says that instead.
Updated, 5 October 2026. A community benchmark of a 32 GB Mac mini M6 (LM Studio, median of three runs, raw data published) measured Qwen 3.6-35B-A3B at 64 tok/s, Gemma 4 26B-A4B at 48, Nemotron 3.5 Lightning at 45 and GPT-oss 20B at 45 — three to four times the mixture-of-experts estimates below. Dense models landed where this article put them: Gemma 4 31B at 8, Muse Glimmer 30B at 9. The index’s mixture-of-experts model has been refit to these runs; current figures are on the Mac mini M6 32 GB page.
Updated 23 September 2026: the M6 shipped on 22 September. No press review had published local-model figures for it as of 23 September; an early hands-on video reports Qwen3.8-27B at 8–9 tok/s with plain MLX on a 32 GB model, where the index’s estimate of 8 sits. How every estimate held up →
Apple’s announcement leads with the 2nm process and the doubled Neural Engine. For local LLMs, neither is the number that matters. Token generation is bandwidth-bound, and the M6’s story is 153 or 170 GB/s depending on which configuration you buy — a distinction Apple prints in the spec table and nowhere else.
The bin trap: 153 vs 170
Per Apple’s spec page, the 16 GB-based configurations run 153 GB/s; the model that starts at 24 GB runs 170. That is a 10% speed difference before you consider memory itself — and memory is the harder limit. macOS lets the GPU address roughly 75% of unified RAM:
Config
Bandwidth
GPU budget
What that means
16 GB ($899)
153 GB/s
~12 GB
7–9B models at 4-bit, nothing more
24 GB (~$1,199)
170 GB/s
~18 GB
14B comfortable; 27B fits with zero headroom
32 GB (~$1,399)
170 GB/s
~24 GB
27B-class at 4-bit with real headroom
Estimated speeds
Model
Needs
M6 est. tok/s
LFM2.5-2.6B
16 GB ok
~81
Gemma 4 E4B
16 GB ok
~55
DeepSeek R1 8B
16 GB ok
~29
Qwen 3.5 9B
16 GB ok
~26
Phi-4 14B
16 GB tight
~17
Gemma 4 26B-A4B (MoE)
24 GB
~14
Nemotron 3.5 Lightning (MoE)
32 GB
~17
Qwen 3.6-35B-A3B (MoE)
32 GB
~14
Qwen3.8-27B (top dense model)
24 GB zero headroom; 32 GB right
~8
Muse Glimmer 30B
32 GB
~8
All estimated at 170 GB/s via the LLMCheck model; subtract ~10% for the 16 GB bins. The machine shipped on 22 September; no published measurement existed as of 23 September.
The honest framing: 8–9 tok/s for a 27B is reading-speed, not conversation-speed. Where the M6 shines is the small-and-sharp tier — MoE models with small active sets (Gemma’s 26B-A4B at ~14, the 35B-A3B coder at ~14) and the sub-10B dense models in the 25–80 range.
$899 mini vs the alternatives
vs Mac mini M4 (discounted): the M4 runs 120 GB/s; the M6’s 170 is a 42% bandwidth jump at the 24 GB tier. The M6 is the better buy unless the M4 drops well under $500.
vs MacBook Air M4 24GB: same 120 GB/s class as the old mini — the M6 mini is now clearly faster for LLM work; the Air’s advantage is the screen and battery, not the inference.
vs Mac mini M5 Pro: $800 more buys 307 GB/s and up to 64 GB — roughly 1.8× the speed and a whole model class. If local AI is the point of the purchase, that is where the money goes.
Buying advice in one line: skip 16 GB, buy 24 GB if the budget stops at ~$1,200, buy 32 GB if you want the 27B-class door open. Every ranked model for each config: 16 GB · 24 GB · 32 GB.
LLMCheck
An independent index of local-LLM performance on Apple Silicon. Every figure is labelled sourced, estimated or community — see the methodology.
Frequently Asked Questions
Can the $899 Mac mini M6 run local LLMs?
Yes, within limits: its 16 GB and 153 GB/s handle 7–9B models at 4-bit around 23–26 tok/s estimated. For anything larger, buy the 24 or 32 GB config (170 GB/s).
Can the Mac mini M6 run Qwen3.8-27B?
At 32 GB, yes — the 4-bit build (~16 GB of weights) fits the ~24 GB GPU budget with headroom, at an estimated ~8 tok/s. The 24 GB config technically fits it with almost no headroom; the 16 GB config cannot load it.
Is the M6 mini faster than the M4 mini for AI?
Meaningfully: 170 vs 120 GB/s is ~42% more decode bandwidth at the 24 GB tier, plus a doubled Neural Engine that helps prompt processing. The 16 GB M6 bins (153 GB/s) still beat the M4 by ~28%.
M6 mini or M5 Pro mini for local AI?
If local AI is the reason for the purchase, the M5 Pro: 307 GB/s and up to 64 GB is a different class (~1.8× faster, runs 70B-adjacent MoE). The M6 is the right buy at $1,400 and below.
🛒 Where to buy
The new machines have been shipping since 22 September:
As an Amazon Associate, LLMCheck earns from qualifying purchases. The links above are affiliate links — they cost you nothing extra and help keep the index free and ad-light.