The M5 Ultra is built from four dies where every previous Ultra fused two, with over 4.4 TB/s of inter-die bandwidth holding it together. For local inference the external number is the one that pays: 1.2 TB/s of unified memory bandwidth, against 819 GB/s for the M3 Ultra and 1,092 for the M4 Ultra.

Estimated performance

ModelMin configM5 Ultra est.vs M3 Ultra 512GB
DeepSeek V4 Flash (284B-A13B)256 GB~80 tok/s~53
Mistral Small 4 (119B MoE)96 GB~86~57
Qwen3-Coder-Next (80B-A3B)96 GB~72~48
GLM-4.5-Air (106B MoE)96 GB~61~41
Qwen3-235B-A22B256 GB~37~25
Inkling-Small (276B MoE)256 GB~45~30
Llama 3.3 70B (dense)96 GB~24~16
Qwen 2.5 72B (dense)96 GB~23~16
Llama 3.1 405B (dense)512 GB~4~3

Everything estimated — the machine ships 22 September and no one has measured one. Dense figures come from the bandwidth formula; MoE figures scale each model’s existing reference figure by the bandwidth ratio; the M3 Ultra column scales the same way at 819 GB/s.

What 512 GB does — and does not — change

The M3 Ultra already offered 512 GB, so the class of model a Mac can hold is unchanged: GLM 5.2 fits, DeepSeek V4 Pro (1.6T) still does not, and a 405B dense model runs at single-digit speeds that make it a demonstration, not a tool. What changes is that everything at the ceiling runs ~50% faster than on the M3 Ultra — and the agentic MoE class (V4 Flash, Coder-Next, GLM-Air) crosses from “usable” into genuinely fast, at 60–80 tok/s estimated.

The config that matters is 256 GB. 96 GB cannot hold the models that justify an Ultra over an M5 Max Studio ($3,000 less). 512 GB buys headroom whose main occupants run too slowly to matter. 256 GB (~$7,099) holds DeepSeek V4 Flash at 4-bit — the strongest agent that fits any Mac — at ~80 tok/s estimated, plus the 235B Qwen and Inkling-Small. That is the sweet spot.

Against the alternatives

Full rankings for each config: 96 GB · 256 GB · 512 GB.