M5 Ultra Mac Studio: 1.2 TB/s and 512 GB, Examined for Local LLMs
The M5 Ultra (25 Aug 2026, from $5,499) is Apple’s first quad-die chip: 1.2 TB/s of memory bandwidth — 50% above M3 Ultra — with 96, 256 or 512 GB of unified memory. Estimated: DeepSeek V4 Flash ~47 tok/s, Qwen3-Coder-Next ~72, a 70B dense ~24. It does not raise the size ceiling; it makes the ceiling fast.
Updated 7 September 2026: the M5 Ultra estimate for DeepSeek V4 Flash was recalibrated from ~80 to ~47 tok/s. The earlier figure applied the dense bandwidth model to a sparse model; the index now scales MoE estimates by a factor fitted to its one community-measured MoE run. Method on /methodology.
The M5 Ultra is built from four dies where every previous Ultra fused two, with over 4.4 TB/s of inter-die bandwidth holding it together. For local inference the external number is the one that pays: 1.2 TB/s of unified memory bandwidth, against 819 GB/s for the M3 Ultra and 1,092 for the M4 Ultra.
Estimated performance
Model
Min config
M5 Ultra est.
vs M3 Ultra 512GB
DeepSeek V4 Flash (284B-A13B)
256 GB
~80 tok/s
~53
Mistral Small 4 (119B MoE)
96 GB
~86
~57
Qwen3-Coder-Next (80B-A3B)
96 GB
~72
~48
GLM-4.5-Air (106B MoE)
96 GB
~61
~41
Qwen3-235B-A22B
256 GB
~37
~25
Inkling-Small (276B MoE)
256 GB
~45
~30
Llama 3.3 70B (dense)
96 GB
~24
~16
Qwen 2.5 72B (dense)
96 GB
~23
~16
Llama 3.1 405B (dense)
512 GB
~4
~3
Everything estimated — the machine ships 22 September and no one has measured one. Dense figures come from the bandwidth formula; MoE figures scale each model’s existing reference figure by the bandwidth ratio; the M3 Ultra column scales the same way at 819 GB/s.
What 512 GB does — and does not — change
The M3 Ultra already offered 512 GB, so the class of model a Mac can hold is unchanged: GLM 5.2 fits, DeepSeek V4 Pro (1.6T) still does not, and a 405B dense model runs at single-digit speeds that make it a demonstration, not a tool. What changes is that everything at the ceiling runs ~50% faster than on the M3 Ultra — and the agentic MoE class (V4 Flash, Coder-Next, GLM-Air) crosses from “usable” into genuinely fast, at 60–80 tok/s estimated.
The config that matters is 256 GB. 96 GB cannot hold the models that justify an Ultra over an M5 Max Studio ($3,000 less). 512 GB buys headroom whose main occupants run too slowly to matter. 256 GB (~$7,099) holds DeepSeek V4 Flash at 4-bit — the strongest agent that fits any Mac — at ~47 tok/s estimated, plus the 235B Qwen and Inkling-Small. That is the sweet spot.
Against the alternatives
M4 Ultra 192 GB (discounted): 1,092 GB/s — 89% of the speed. What it lacks is the 256 GB tier, which is exactly where V4 Flash at 4-bit lives. If you do not need that model, the discount wins.
Used M3 Ultra 512 GB: the budget half-terabyte. Same ceiling, two-thirds the speed, likely half the price on the used market from today.
M5 Max Studio 128 GB: $2,499–$3,999 covers everything through 70B dense and V4 Flash at 2-bit. The Ultra is for the 4-bit-big-MoE tier and nothing smaller justifies it.
An independent index of local-LLM performance on Apple Silicon. Every figure is labelled sourced, estimated or community — see the methodology.
Frequently Asked Questions
How fast is the M5 Ultra for local LLMs?
Estimated from Apple’s published 1.2 TB/s: DeepSeek V4 Flash ~47 tok/s, Qwen3-Coder-Next ~72, GLM-4.5-Air ~61, 70B dense ~24. Nothing is measured yet — it ships 22 September.
Which M5 Ultra memory config should I buy for AI?
256 GB. 96 GB cannot hold the models that justify the Ultra’s price over an M5 Max; 512 GB mostly buys room for models too slow to use. 256 GB holds DeepSeek V4 Flash at 4-bit — the strongest Mac-runnable agent — at speed.
Can the M5 Ultra run GLM 5.2 or Kimi K3?
GLM 5.2 (256 GB minimum) fits the 512 GB config, as it did on M3 Ultra; no per-machine speed figure is published for it. Kimi K3 (~1.5 TB class) and DeepSeek V4 Pro remain server-only.
Is the M5 Ultra worth it over a discounted M4 Ultra?
Only if you need the 256 GB tier — that is where DeepSeek V4 Flash at 4-bit lives, and the M4 Ultra tops out at 192 GB. Otherwise the M4 Ultra keeps 89% of the bandwidth at a growing discount.
🛒 Where to buy
Preorders are Apple-only until units ship on 22 September — Amazon has no new-machine listings yet, so until then these buttons land on the closest currently-buyable configs (largely the discounted outgoing generation):
As an Amazon Associate, LLMCheck earns from qualifying purchases. The links above are affiliate links — they cost you nothing extra and help keep the index free and ad-light.