Apple’s announcement leads with the 2nm process and the doubled Neural Engine. For local LLMs, neither is the number that matters. Token generation is bandwidth-bound, and the M6’s story is 153 or 170 GB/s depending on which configuration you buy — a distinction Apple prints in the spec table and nowhere else.

The bin trap: 153 vs 170

Per Apple’s spec page, the 16 GB-based configurations run 153 GB/s; the model that starts at 24 GB runs 170. That is a 10% speed difference before you consider memory itself — and memory is the harder limit. macOS lets the GPU address roughly 75% of unified RAM:

ConfigBandwidthGPU budgetWhat that means
16 GB ($899)153 GB/s~12 GB7–9B models at 4-bit, nothing more
24 GB (~$1,199)170 GB/s~18 GB14B comfortable; 27B fits with zero headroom
32 GB (~$1,399)170 GB/s~24 GB27B-class at 4-bit with real headroom

Estimated speeds

ModelNeedsM6 est. tok/s
LFM2.5-2.6B16 GB ok~81
Gemma 4 E4B16 GB ok~55
DeepSeek R1 8B16 GB ok~29
Qwen 3.5 9B16 GB ok~26
Phi-4 14B16 GB tight~17
Gemma 4 26B-A4B (MoE)24 GB~14
Nemotron 3.5 Lightning (MoE)32 GB~17
Qwen 3.6-35B-A3B (MoE)32 GB~14
Qwen3.8-27B (index #1)24 GB zero headroom; 32 GB right~8
Muse Glimmer 30B32 GB~8

All estimated at 170 GB/s via the LLMCheck model; subtract ~10% for the 16 GB bins. The machine ships 22 September — no measured figures exist yet from anyone.

The honest framing: 8–9 tok/s for a 27B is reading-speed, not conversation-speed. Where the M6 shines is the small-and-sharp tier — MoE models with small active sets (Gemma’s 26B-A4B at ~14, the 35B-A3B coder at ~14) and the sub-10B dense models in the 25–80 range.

$899 mini vs the alternatives

Buying advice in one line: skip 16 GB, buy 24 GB if the budget stops at ~$1,200, buy 32 GB if you want the 27B-class door open. Every ranked model for each config: 16 GB · 24 GB · 32 GB.