M5 Ultra Mac Studio: 1.2 TB/s and 512 GB, Examined for Local LLMs
The M5 Ultra (25 Aug 2026, from $5,499) is Apple’s first quad-die chip: 1.2 TB/s of memory bandwidth — 50% above M3 Ultra — with 96, 256 or 512 GB of unified memory. Estimated: DeepSeek V4 Flash ~47 tok/s, Qwen3-Coder-Next ~72, a 70B dense ~24. It does not raise the size ceiling; it makes the ceiling fast.
Correction, 5 October 2026. Earlier versions of this article cited a Mac Studio with an “M4 Ultra” chip. Apple never released one — the M4 Max has no UltraFusion connector, and the 2025 Mac Studio paired it with the M3 Ultra (819 GB/s, up to 512 GB). The comparisons now use Macs that exist, with figures from the current LLMCheck estimator.
Updated, 5 October 2026. Since launch, the index’s mixture-of-experts estimates were refit to the first fully specified measurements of small MoE models, which ran two to four times faster than the flat factor the index had used. Current figures for every Mac are on the Best-by-Mac pages and the method is on /methodology. On an M5 Ultra the index now puts Qwen3-Coder-Next at about 130 tok/s (this article said ~72), GLM-4.5-Air at about 49 (~61) and DeepSeek V4 Flash at about 45 (~47); dense models are unchanged.
Updated 23 September 2026: the M5 Ultra shipped on 22 September and the first reviews are in. BGR measured Qwen3.8-27B at about 55 tok/s on an M5 Ultra and 40 on an M3 Ultra — within 4% of the index’s estimates — and Qwen 3.5 122B-A10B at about 80 tok/s, well above what the index’s conservative MoE factor predicts. MacStories measured prompt processing about 2.5× faster than the M3 Ultra. How every estimate held up →
Updated 7 September 2026: the M5 Ultra estimate for DeepSeek V4 Flash was recalibrated from ~80 to ~47 tok/s. The earlier figure applied the dense bandwidth model to a sparse model; the index now scales MoE estimates by a factor fitted to its one community-measured MoE run. Method on /methodology.
The M5 Ultra is built from four dies where every previous Ultra fused two, with over 4.4 TB/s of inter-die bandwidth holding it together. For local inference the external number is the one that pays: 1.2 TB/s of unified memory bandwidth, against 819 GB/s for the M3 Ultra.
Estimated performance
Model
Min config
M5 Ultra est.
vs M3 Ultra 512GB
DeepSeek V4 Flash (284B-A13B)
256 GB
~80 tok/s
~53
Mistral Small 4 (119B MoE)
96 GB
~86
~57
Qwen3-Coder-Next (80B-A3B)
96 GB
~72
~48
GLM-4.5-Air (106B MoE)
96 GB
~61
~41
Qwen3-235B-A22B
256 GB
~37
~25
Inkling-Small (276B MoE)
256 GB
~45
~30
Llama 3.3 70B (dense)
96 GB
~24
~16
Qwen 2.5 72B (dense)
96 GB
~23
~16
Llama 3.1 405B (dense)
512 GB
~4
~3
Everything below was estimated before the machine shipped on 22 September; the first measurements are in the note at the top. Dense figures come from the bandwidth formula; MoE figures scale each model’s existing reference figure by the bandwidth ratio; the M3 Ultra column scales the same way at 819 GB/s.
What 512 GB does — and does not — change
The M3 Ultra already offered 512 GB, so the class of model a Mac can hold is unchanged: GLM 5.2 fits, DeepSeek V4 Pro (1.6T) still does not, and a 405B dense model runs at single-digit speeds that make it a demonstration, not a tool. What changes is that everything at the ceiling runs ~50% faster than on the M3 Ultra — and the agentic MoE class (V4 Flash, Coder-Next, GLM-Air) crosses from “usable” into genuinely fast, at 60–80 tok/s estimated.
The config that matters is 256 GB. 96 GB cannot hold the models that justify an Ultra over an M5 Max Studio ($3,000 less). 512 GB buys headroom whose main occupants run too slowly to matter. 256 GB (~$7,099) holds DeepSeek V4 Flash at 4-bit — the strongest agent that fits any Mac — at ~47 tok/s estimated, plus the 235B Qwen and Inkling-Small. That is the sweet spot.
Against the alternatives
Used M3 Ultra 512 GB: the budget half-terabyte. Same ceiling, two-thirds the speed, likely half the price on the used market from today.
M5 Max Studio 128 GB: $2,499–$3,999 covers everything through 70B dense and V4 Flash at 2-bit. The Ultra is for the 4-bit-big-MoE tier and nothing smaller justifies it.
An independent index of local-LLM performance on Apple Silicon. Every figure is labelled sourced, estimated or community — see the methodology.
Frequently Asked Questions
How fast is the M5 Ultra for local LLMs?
Estimated from Apple’s published 1.2 TB/s: DeepSeek V4 Flash ~47 tok/s, Qwen3-Coder-Next ~72, GLM-4.5-Air ~61, 70B dense ~24. Nothing is measured yet — it ships 22 September.
Which M5 Ultra memory config should I buy for AI?
256 GB. 96 GB cannot hold the models that justify the Ultra’s price over an M5 Max; 512 GB mostly buys room for models too slow to use. 256 GB holds DeepSeek V4 Flash at 4-bit — the strongest Mac-runnable agent — at speed.
Can the M5 Ultra run GLM 5.2 or Kimi K3?
GLM 5.2 (256 GB minimum) fits the 512 GB config, as it did on M3 Ultra; no per-machine speed figure is published for it. Kimi K3 (~1.5 TB class) and DeepSeek V4 Pro remain server-only.
Is the M5 Ultra worth it over a discounted M3 Ultra?
For speed, yes: 1.2 TB/s is 50% more bandwidth than the M3 Ultra’s 819 GB/s, and BGR measured Qwen3.8-27B at about 55 tok/s against 40 on the M3 Ultra. For capacity alone, a discounted M3 Ultra with 256 or 512 GB holds the same models. Apple never made an M4 Ultra.
🛒 Where to buy
The new machines have been shipping since 22 September:
As an Amazon Associate, LLMCheck earns from qualifying purchases. The links above are affiliate links — they cost you nothing extra and help keep the index free and ad-light.