Updated August 2026 · 202 Benchmarks

Apple Silicon LLM Benchmarks

According to LLMCheck, this open dataset covers 248 benchmark data points across 61 local LLMs and 16 Apple Silicon chips — real tokens-per-second and time-to-first-token for each. Every figure is labelled with its provenance: 5 vendor-published, 5 linked third-party runs, and 217 derived from the published memory-bandwidth model at Q4_K_M. Free to download as CSV or JSON.

🔍 Not sure what your Mac can run? Check your Mac in 10 seconds →

ⓘ Figures are transparent estimates unless marked sourced/community. Own a Mac? Submit a real benchmark →

Tokens-per-second figures — estimated, sourced, and community-submitted, each row labeled — across 56 models, 16 Apple Silicon chips, and 3 inference engines. Find exactly how fast your model runs on your Mac.

56
Models
202
Data Points
14
Chips Tested
3
Engines
Chip:
RAM:
Engine:
Model Params Quant Chip RAM Engine tok/s (est.) TTFT Date
Phi-4 Mini3.8BQ4_K_MM5 Max64 GBOllamaestimated1680.3s2026-03
Phi-4 Mini3.8BQ4_K_MM4 Pro24 GBOllamaestimated880.4s2026-03
Phi-4 Mini3.8BQ4_K_MM316 GBMLXestimated350.3s2026-02
Phi-4 Mini3.8BQ4_K_MM28 GBOllamaestimated350.5s2026-01
Phi-4 Mini3.8BQ4_K_MM116 GBOllamaestimated240.6s2025-12
Qwen 3.5 4B4BQ4_K_MM5 Max64 GBMLXestimated1610.2s2026-03
Qwen 3.5 4B4BQ4_K_MM416 GBOllamaestimated400.4s2026-02
Qwen 3 4B4BQ4_K_MM28 GBOllamaestimated330.7s2026-01
Qwen 3 4B4BQ4_K_MM4 Pro24 GBMLXestimated840.3s2026-02
Gemma 4 E2B2.3BQ4_K_MM5 Max128 GBMLXestimated2610.1s2026-04
Gemma 4 E2B2.3BQ4_K_MM4 Pro24 GBOllamaestimated1500.2s2026-04
Gemma 4 E2B2.3BQ4_K_MM316 GBOllamaestimated640.3s2026-04
Gemma 4 E2B2.3BQ4_K_MM18 GBOllamaestimated450.5s2026-04
Gemma 4 E4B4BQ4_K_MM5 Max128 GBMLXestimated1610.2s2026-04
Gemma 4 E4B4BQ4_K_MM5 Pro24 GBOllamaestimated840.3s2026-04
Gemma 4 E4B4BQ4_K_MM4 Pro24 GBMLXestimated840.3s2026-04
Gemma 4 E4B4BQ4_K_MM316 GBOllamaestimated330.4s2026-04
Gemma 4 E4B4BQ4_K_MM18 GBOllamaestimated230.6s2026-04
Gemma 4 26B-A4B26BQ4_K_MM5 Max128 GBMLXestimated500.5s2026-04
Gemma 4 26B-A4B26BQ4_K_MM5 Pro24 GBOllamaestimated350.8s2026-04
Gemma 4 26B-A4B26BQ4_K_MM4 Max48 GBMLXestimated400.7s2026-04
Gemma 4 26B-A4B26BQ4_K_MM4 Pro24 GBOllamaestimated281.0s2026-04
Gemma 4 31B31BQ4_K_MM5 Max128 GBMLXestimated260.7s2026-04
Gemma 4 31B31BQ4_K_MM5 Max64 GBOllamaestimated260.9s2026-04
Gemma 4 31B31BQ4_K_MM4 Max48 GBMLXestimated241.1s2026-04
Gemma 4 31B31BQ4_K_MM4 Pro24 GBOllamaestimated121.4s2026-04
Gemma 3 4B4BQ4_K_MM5 Max64 GBOllamaestimated1610.3s2026-03
Gemma 3 4B4BQ4_K_MM18 GBOllamaestimated230.8s2025-12
Gemma 3 4B4BQ4_K_MM3 Pro18 GBMLXestimated490.4s2026-01
Qwen 3.5 9B9BQ4_K_MM5 Max64 GBOllamaestimated820.5s2026-03
Qwen 3.5 9B9BQ4_K_MM416 GBLM Studioestimated180.6s2026-02
Qwen 3.5 9B9BQ4_K_MM316 GBOllamaestimated150.7s2026-01
Qwen 3.5 9B9BQ4_K_MM116 GBOllamaestimated101.1s2025-12
Qwen 3 8B8BQ4_K_MM5 Max128 GBOllamaestimated910.4s2026-03
Qwen 3 8B8BQ8_0M4 Pro24 GBMLXestimated450.5s2026-02
Qwen 3 8B8BQ4_K_MM216 GBOllamaestimated170.7s2026-01
DeepSeek R1 8B8BQ4_K_MM5 Max64 GBOllamaestimated910.5s2026-03
DeepSeek R1 8B8BQ4_K_MM416 GBMLXestimated200.5s2026-02
DeepSeek R1 8B8BQ4_K_MM216 GBOllamaestimated170.8s2026-01
DeepSeek R1 8B8BQ4_K_MM116 GBOllamaestimated121.2s2025-11
Mistral 7B7BQ4_K_MM5 Max64 GBOllamaestimated1020.3s2026-03
Mistral 7B7BQ4_K_MM4 Pro24 GBMLXestimated510.4s2026-02
Mistral 7B7BQ4_K_MM316 GBOllamaestimated200.6s2026-01
Mistral 7B7BQ4_K_MM116 GBOllamaestimated131.2s2025-12
Llama 3.1 8B8BQ4_K_MM5 Max128 GBMLXestimated910.3s2026-03
Llama 3.1 8B8BQ4_K_MM416 GBOllamaestimated200.6s2026-02
Llama 3.1 8B8BQ4_K_MM3 Pro18 GBOllamaestimated250.7s2026-01
Llama 3.1 8B8BQ4_K_MM116 GBOllamaestimated121.1s2025-12
Ministral 8B8BQ4_K_MM5 Max64 GBOllamaestimated910.4s2026-03
Ministral 8B8BQ4_K_MM316 GBLM Studioestimated170.7s2026-01
Gemma 3 12B12BQ4_K_MM5 Max64 GBOllamaestimated640.6s2026-03
Gemma 3 12B12BQ4_K_MM4 Pro24 GBMLXestimated310.7s2026-02
Gemma 3 12B12BQ4_K_MM316 GBOllamaestimated121.1s2026-01
Gemma 3 12B12BQ4_K_MM216 GBOllamaestimated121.3s2025-12
Qwen 3 14B14BQ4_K_MM5 Max64 GBOllamaestimated550.8s2026-03
Qwen 3 14B14BQ4_K_MM4 Pro24 GBLM Studioestimated261.0s2026-02
Qwen 3 14B14BQ4_K_MM316 GBOllamaestimated101.2s2026-01
Phi-4 14B14BQ4_K_MM5 Max64 GBMLXestimated550.6s2026-03
Phi-4 14B14BQ4_K_MM416 GBOllamaestimated121.0s2026-02
Phi-4 14B14BQ4_K_MM216 GBMLXestimated101.3s2026-01
Ministral 3 14B14BQ4_K_MM5 Max64 GBOllamaestimated550.7s2026-03
Ministral 3 14B14BQ4_K_MM4 Pro24 GBOllamaestimated260.9s2026-02
Gemma 3 27B27BQ4_K_MM5 Max64 GBOllamaestimated300.9s2026-03
Gemma 3 27B27BQ4_K_MM4 Pro24 GBLM Studioestimated141.5s2026-02
Gemma 3 27B27BQ4_K_MM4 Max48 GBMLXestimated271.1s2026-02
Qwen 3 30B-A3B30BQ4_K_MM5 Max64 GBOllamaestimated620.7s2026-03
Qwen 3 30B-A3B30BQ4_K_MM4 Max48 GBOllamaestimated420.9s2026-02
Qwen 3 30B-A3B30BQ4_K_MM4 Pro24 GBMLXestimated351.0s2026-02
Qwen 3 30B-A3B30BQ4_K_MM3 Max36 GBOllamaestimated411.3s2026-01
Qwen 3.5 35B-A3B35BQ4_K_MM5 Max64 GBMLXestimated520.9s2026-03
Qwen 3.5 35B-A3B35BQ4_K_MM4 Max48 GBOllamaestimated341.2s2026-02
Qwen 3.5 35B-A3B35BQ4_K_MM5 Max128 GBOllamaestimated480.9s2026-03
Qwen 3 32B32BQ4_K_MM5 Max128 GBOllamaestimated251.1s2026-03
Qwen 3 32B32BQ4_K_MM4 Max64 GBMLXestimated231.4s2026-02
Qwen 3 32B32BQ4_K_MM4 Pro32 GBOllamaestimated121.8s2026-01
DeepSeek R1 32B32BQ4_K_MM5 Max64 GBOllamaestimated251.2s2026-03
DeepSeek R1 32B32BQ4_K_MM4 Max48 GBLM Studioestimated231.8s2026-01
DeepSeek R1 32B32BQ4_K_MM3 Max36 GBOllamaestimated172.0s2025-12
Llama 3.3 70B70BQ4_K_MM5 Max128 GBOllamaestimated122.8s2026-03
Llama 3.3 70B70BQ4_K_MM4 Ultra192 GBMLXestimated212.0s2026-02
Llama 3.3 70B70BQ4_K_MM5 Max128 GBMLXestimated122.4s2026-03
DeepSeek R1 70B70BQ4_K_MM5 Max128 GBOllamaestimated123.0s2026-03
DeepSeek R1 70B70BQ4_K_MM4 Ultra192 GBMLXestimated212.2s2026-02
Qwen 2.5 72B72BQ4_K_MM5 Max128 GBOllamaestimated123.2s2026-03
Qwen 2.5 72B72BQ4_K_MM4 Ultra192 GBOllamaestimated212.5s2026-02
GPT-oss 120B120BQ4_K_MM5 Max128 GBOllamaestimated75.5s2026-03
GPT-oss 120B120BQ4_K_MM4 Ultra192 GBMLXestimated134.2s2026-02
Llama 4 Scout109BQ4_K_MM5 Max128 GBOllamaestimated221.8s2026-03
Llama 4 Scout109BQ4_K_MM4 Ultra192 GBMLXestimated301.3s2026-02
Llama 4 Scout109BQ4_K_MM5 Max128 GBMLXestimated261.5s2026-03
Mistral 7B7BQ8_0M5 Max64 GBMLXestimated1020.4s2026-03
Mistral 7B7BQ8_0M4 Pro24 GBOllamaestimated510.5s2026-02
Qwen 3.5 9B9BQ8_0M5 Max128 GBMLXestimated820.5s2026-03
Llama 3.1 8B8BQ8_0M5 Max64 GBOllamaestimated910.4s2026-03
Qwen 3 14B14BQ8_0M5 Max128 GBMLXestimated550.9s2026-03
Phi-4 Mini3.8BQ8_0M5 Max64 GBMLXestimated1680.3s2026-03
Phi-4 Mini3.8BQ4_K_MM4 Max48 GBMLXestimated1560.3s2026-02
Gemma 3 4B4BQ8_0M4 Pro24 GBMLXestimated840.3s2026-02
Qwen 3.5 35B-A3B35BQ4_K_MM3 Max96 GBOllamaestimated221.5s2026-01
DeepSeek R1 8B8BQ8_0M5 Max64 GBMLXestimated910.5s2026-03
Qwen 3 8B8BQ4_K_MM4 Pro24 GBOllamaestimated450.5s2026-02
Qwen 3 8B8BQ4_K_MM116 GBOllamaestimated121.0s2025-12
Ministral 8B8BQ4_K_MM416 GBMLXestimated200.5s2026-02
Ministral 3 14B14BQ4_K_MM316 GBLM Studioestimated101.2s2026-01
Gemma 3 27B27BQ4_K_MM3 Max96 GBMLXestimated201.3s2026-01
Qwen 3 32B32BQ4_K_MM4 Ultra192 GBOllamaestimated451.0s2026-02
Llama 3.1 8B8BQ4_K_MM28 GBOllamaestimated170.8s2025-12
Qwen 3.5 9B9BQ4_K_MM4 Pro24 GBMLXestimated400.4s2026-03
Qwen 3 4B4BQ4_K_MM5 Max64 GBOllamaestimated1610.2s2026-03
Qwen 3.6-35B-A3B35BQ4_K_MM5 Max128 GBMLXestimated550.6s2026-04
Qwen 3.6-35B-A3B35BQ4_K_MM5 Max64 GBOllamaestimated480.8s2026-04
Qwen 3.6-35B-A3B35BQ4_K_MM4 Max48 GBMLXestimated420.9s2026-04
Qwen 3.6-35B-A3B35BQ4_K_MM4 Pro24 GBOllamaestimated321.2s2026-04
Mistral Small 4119BQ4_K_MM5 Max128 GBOllamaestimated380.9s2026-04
Mistral Small 4119BQ4_K_MM5 Max128 GBMLXestimated420.8s2026-04
Mistral Small 4119BQ4_K_MM4 Ultra192 GBMLXestimated450.7s2026-04
Qwen3-235B-A22B235BQ4_K_MM5 Max128 GBOllamaestimated152.2s2026-04
Qwen3-235B-A22B235BQ4_K_MM5 Max128 GBMLXestimated181.8s2026-04
Qwen3-235B-A22B235BQ4_K_MM4 Ultra192 GBMLXestimated221.5s2026-04
Nemotron-Cascade 230BQ4_K_MM5 Max64 GBOllamaestimated350.9s2026-04
Nemotron-Cascade 230BQ4_K_MM4 Max48 GBMLXestimated281.2s2026-04
Nemotron-Cascade 230BQ4_K_MM4 Pro24 GBOllamaestimated221.5s2026-04
Mistral Small 3.2 24B24BQ4_K_MM5 Max128 GBMLXestimated330.6s2026-05
Mistral Small 3.2 24B24BQ4_K_MM5 Max64 GBOllamaestimated330.7s2026-05
Mistral Small 3.2 24B24BQ4_K_MM4 Max48 GBMLXestimated310.8s2026-05
Mistral Small 3.2 24B24BQ4_K_MM4 Pro32 GBOllamaestimated161.0s2026-05
Hermes 4 70B70BQ4_K_MM5 Max128 GBMLXestimated121.8s2026-05
Hermes 4 70B70BQ4_K_MM4 Max128 GBMLXestimated112.1s2026-05
Hermes 4 70B70BQ4_K_MM4 Ultra192 GBOllamaestimated211.5s2026-05
SmolLM3 3B3BQ4_K_MM5 Max64 GBMLXestimated1990.1s2026-05
SmolLM3 3B3BQ4_K_MM4 Pro24 GBOllamaestimated1080.2s2026-05
SmolLM3 3B3BQ4_K_MM316 GBOllamaestimated440.3s2026-05
SmolLM3 3B3BQ4_K_MM28 GBOllamaestimated440.4s2026-05
SmolLM3 3B3BQ4_K_MM18 GBOllamaestimated300.5s2026-05
Devstral Small 24B24BQ4_K_MM5 Max128 GBMLXestimated330.6s2026-05
Devstral Small 24B24BQ4_K_MM5 Max64 GBOllamaestimated330.7s2026-05
Devstral Small 24B24BQ4_K_MM4 Max48 GBMLXestimated310.8s2026-05
Devstral Small 24B24BQ4_K_MM4 Pro32 GBOllamaestimated161.0s2026-05
GLM-4.5-Air106BQ4_K_MM5 Max128 GBMLXestimated340.9s2026-07
GLM-4.5-Air106BQ4_K_MM5 Max64 GBOllamaestimated301.1s2026-07
GLM-4.5-Air106BQ4_K_MM4 Max128 GBMLXestimated261.3s2026-07
GLM-4.5-Air106BQ4_K_MM4 Ultra192 GBMLXestimated380.8s2026-07
DeepSeek V4 Flash284B-A13BQ2_KM5 Max128 GBMLXcommunity390.4s2026-08
DeepSeek V4 Flash284B-A13BQ2_KM3 Max128 GBMLXcommunity270.6s2026-08
GLM 5.2753B-A40BIQ1_SM3 Ultra256 GBLM Studiocommunity221.5s2026-07
GLM 5.2753B-A40BQ4_K_MM3 Ultra512 GBMLXcommunity152.0s2026-07
Muse Glimmer 30B30BQ4_K_MM4 Pro24 GBLM Studiocommunity100.8s2026-08
Muse Glimmer 30B30BQ4_K_MM5 Max64 GBMLXsourced270.4s2026-08
Muse Glimmer 30B30BQ4_K_MM4 Max48 GBMLXsourced240.5s2026-08
LFM2.5-2.6B2.6BQ4_K_MM5 Max64 GBMLXsourced2200.1s2026-08
Maple Preview 20B-A1B20B-A1BTernaryM5 Pro24 GBMLXsourced2810.1s2026-08
Maple Preview 20B-A1B20B-A1BTernaryM416 GBMLXsourced2000.1s2026-08
Qwen 3.6-27B27BQ4_K_MM5 Max64 GBMLXestimated300.3s2026-08
Qwen 3.6-27B27BQ4_K_MM4 Max64 GBMLXestimated270.3s2026-08
Qwen 3.6-27B27BQ4_K_MM3 Max64 GBOllamaestimated200.4s2026-08
Qwen 3.6-27B27BQ4_K_MM5 Pro24 GBMLXestimated140.5s2026-08
Qwen 3.6-27B27BQ4_K_MM4 Pro24 GBOllamaestimated140.5s2026-08
KAT-Coder-V2.535B-A3BQ4_K_MM5 Max64 GBMLXestimated520.3s2026-08
KAT-Coder-V2.535B-A3BQ4_K_MM4 Max48 GBMLXestimated470.3s2026-08
KAT-Coder-V2.535B-A3BQ4_K_MM4 Pro24 GBOllamaestimated240.4s2026-08
Nemotron 3.5 Lightning30B-A3BQ4_K_MM5 Max64 GBMLXestimated600.2s2026-08
Nemotron 3.5 Lightning30B-A3BQ4_K_MM4 Max48 GBMLXestimated550.3s2026-08
Nemotron 3.5 Lightning30B-A3BQ4_K_MM4 Pro24 GBOllamaestimated280.4s2026-08
Laguna S 2.1118B-A8BQ4_K_MM5 Max128 GBMLXestimated260.6s2026-08
Laguna S 2.1118B-A8BQ4_K_MM4 Max128 GBMLXestimated240.7s2026-08
Laguna XS 2.133B-A3BQ4_K_MM5 Max64 GBMLXestimated550.3s2026-08
Laguna XS 2.133B-A3BQ4_K_MM4 Pro24 GBOllamaestimated250.4s2026-08
Inkling-Small276B-A12BQ2_KM5 Max128 GBMLXestimated220.8s2026-08
Bonsai 27B27B1-bitM5 Max64 GBMLXestimated300.2s2026-08
Bonsai 27B27B1-bitM416 GBMLXestimated60.3s2026-08
Bonsai 27B27B1-bitM28 GBMLXestimated50.5s2026-08
Nanbeige4.2-3B3BQ4_K_MM5 Max64 GBMLXestimated1990.2s2026-08
Nanbeige4.2-3B3BQ4_K_MM416 GBOllamaestimated520.3s2026-08
Nanbeige4.2-3B3BQ4_K_MM28 GBOllamaestimated440.4s2026-08
Apertus 1.5 8B8BQ4_K_MM5 Max64 GBMLXestimated910.3s2026-08
Apertus 1.5 8B8BQ4_K_MM416 GBOllamaestimated200.4s2026-08
Apertus 1.5 70B70BQ4_K_MM5 Max128 GBMLXestimated120.9s2026-08
Apertus 1.5 70B70BQ4_K_MM4 Ultra96 GBMLXestimated210.6s2026-08
GLM-4.7-Flash31BQ4_K_MM5 Max64 GBMLXestimated260.3s2026-08
GLM-4.7-Flash31BQ4_K_MM4 Max48 GBMLXestimated240.4s2026-08
GLM-4.7-Flash31BQ4_K_MM4 Pro24 GBOllamaestimated120.5s2026-08
MiniMax M2.5230B-A10BQ4_K_MM4 Ultra192 GBMLXestimated330.7s2026-08
Llama 3.3 70B70BQ4_K_MM1 Ultra128 GBOllamaestimated160.8s2026-08
DeepSeek R1 70B70BQ4_K_MM1 Ultra128 GBOllamaestimated160.9s2026-08
Qwen 3.6-35B-A3B35B-A3BQ4_K_MM1 Ultra64 GBMLXestimated640.3s2026-08
Qwen 3.6-27B27BQ4_K_MM1 Ultra64 GBMLXestimated390.4s2026-08
Gemma 4 26B-A4B26B-A4BQ4_K_MM1 Ultra64 GBOllamaestimated670.3s2026-08
GPT-oss 120B117BQ4_K_MM1 Ultra128 GBOllamaestimated91.0s2026-08
Llama 3.3 70B70BQ4_K_MM2 Max96 GBOllamaestimated81.2s2026-08
Qwen 3.6-35B-A3B35B-A3BQ4_K_MM2 Max64 GBMLXestimated320.5s2026-08
Qwen 3.6-27B27BQ4_K_MM2 Max32 GBMLXestimated200.6s2026-08
Gemma 3 27B27BQ4_K_MM2 Max32 GBOllamaestimated200.6s2026-08
Llama 3.1 8B8BQ4_K_MM2 Max32 GBOllamaestimated640.3s2026-08
Mistral Small 3.2 24B24BQ4_K_MM2 Max32 GBOllamaestimated230.5s2026-08
Llama 3.3 70B70BQ4_K_MM3 Ultra256 GBOllamaestimated160.7s2026-08
DeepSeek V4 Flash284B-A13BQ4_K_MM3 Ultra256 GBMLXestimated300.5s2026-08
Inkling-Small276B-A12BQ4_K_MM3 Ultra512 GBMLXestimated300.9s2026-08
Qwen3-235B-A22B235B-A22BQ4_K_MM3 Ultra256 GBMLXestimated190.8s2026-08
Hunyuan Hy3295B-A21BQ4_K_MM3 Ultra256 GBLM Studioestimated160.9s2026-08
Qwen3-Coder-Next80B-A3BQ4_K_MM5 Max64 GBMLXestimated350.4s2026-08
Qwen3-Coder-Next80B-A3BQ4_K_MM4 Max64 GBMLXestimated320.4s2026-08
Qwen3-Coder-Next80B-A3BQ4_K_MM4 Ultra96 GBMLXestimated580.3s2026-08
Qwen3.8-27B27.8BQ4_K_MM624 GBMLXestimated82026-08
Qwen 3.6-27B27BQ4_K_MM624 GBMLXestimated92026-08
Muse Glimmer 30B30BQ4_K_MM632 GBMLXestimated82026-08
Gemma 4 26B-A4B26BQ4_K_MM624 GBMLXestimated142026-08
KAT-Coder-V2.535BQ4_K_MM632 GBMLXestimated152026-08
Qwen 3.6-35B-A3B35BQ4_K_MM632 GBMLXestimated142026-08
Nemotron 3.5 Lightning30BQ4_K_MM632 GBMLXestimated172026-08
DeepSeek R1 8B8BQ4_K_MM616 GBMLXestimated292026-08
Qwen 3.5 9B9BQ4_K_MM616 GBMLXestimated262026-08
Phi-4 14B14BQ4_K_MM616 GBMLXestimated172026-08
Gemma 4 E4B4BQ4_K_MM616 GBMLXestimated552026-08
LFM2.5-2.6B2.6BQ4_K_MM616 GBMLXestimated812026-08
DeepSeek V4 Flash284B-A13BQ4_K_MM5 Ultra256 GBMLXestimated472026-08
Inkling-Small276BQ4_K_MM5 Ultra256 GBMLXestimated452026-08
Qwen3-235B-A22B235BQ4_K_MM5 Ultra256 GBMLXestimated372026-08
Mistral Small 4119BQ4_K_MM5 Ultra96 GBMLXestimated862026-08
GPT-oss 120B117BQ4_K_MM5 Ultra96 GBMLXestimated142026-08
Llama 3.3 70B70BQ4_K_MM5 Ultra96 GBMLXestimated242026-08
DeepSeek R1 70B70BQ4_K_MM5 Ultra96 GBMLXestimated242026-08
Qwen3-Coder-Next80BQ4_K_MM5 Ultra96 GBMLXestimated722026-08
GLM-4.5-Air106BQ4_K_MM5 Ultra96 GBMLXestimated612026-08
Laguna S 2.1118BQ4_K_MM5 Ultra96 GBMLXestimated532026-08
Llama 3.1 405B405BQ4_K_MM5 Ultra512 GBMLXestimated42026-08
Qwen3.8-27B27.8BQ4_K_MM5 Ultra96 GBMLXestimated572026-08
Qwen 2.5 72B72BQ4_K_MM5 Ultra96 GBMLXestimated232026-08
GLM-5.3-Flash320B-A18BIQ3_XXSM5 Max128 GBMLXestimated272026-09
GLM-5.3-Flash320B-A18BIQ4_XSM4 Ultra192 GBMLXestimated322026-09
GLM-5.3-Flash320B-A18BQ4_K_MM3 Ultra256 GBMLXestimated252026-09
GLM-5.3-Flash320B-A18BQ4_K_MM5 Ultra256 GBMLXestimated352026-09
Qwen3.8-Flash-Next125B-A6BQ4_K_MM5 Max128 GBMLXestimated492026-09
Qwen3.8-Flash-Next125B-A6BQ4_K_MM4 Max128 GBMLXestimated452026-09
Qwen3.8-Flash-Next125B-A6BQ4_K_MM4 Ultra192 GBMLXestimated782026-09
Qwen3.8-Flash-Next125B-A6BQ4_K_MM3 Ultra256 GBMLXestimated632026-09
Qwen3.8-Flash-Next125B-A6BQ4_K_MM5 Ultra256 GBMLXestimated852026-09
DeepSeek V4 Flash Vision-Exp305B-A13BQ2_KM5 Max128 GBllama.cppestimated392026-09
DeepSeek V4 Flash Vision-Exp305B-A13BQ4_K_MM4 Ultra192 GBllama.cppestimated422026-09
DeepSeek V4 Flash Vision-Exp305B-A13BQ4_K_MM3 Ultra256 GBllama.cppestimated332026-09
DeepSeek V4 Flash Vision-Exp305B-A13BQ4_K_MM5 Ultra256 GBllama.cppestimated472026-09
Qwen3.8-27B27.8BQ4_K_MM5 Max128 GBMLXestimated292026-09
Qwen3.8-27B27.8BQ4_K_MM4 Max64 GBMLXestimated272026-09
Qwen3.8-27B27.8BQ4_K_MM3 Max64 GBMLXestimated202026-09
Qwen3.8-27B27.8BQ4_K_MM4 Pro48 GBOllamaestimated142026-09
GLM-5.3753B-A40BQ2_KM5 Ultra256 GBMLXestimated292026-09
GLM-5.3753B-A40BQ2_KM3 Ultra256 GBMLXestimated202026-09
GLM-5.3753B-A40BQ4_K_MM5 Ultra512 GBMLXestimated172026-09
GLM-5.3753B-A40BQ4_K_MM3 Ultra512 GBMLXestimated122026-09

Methodology

According to the LLMCheck index, all benchmarks measure tokens per second (tok/s) during the generation phase, excluding prompt processing time. This reflects the sustained output speed you experience when the model is actively generating text.

Time to first token (TTFT) is measured separately in seconds — the delay between submitting your prompt and receiving the first output token. TTFT depends on prompt length, model size, and available memory bandwidth.

Unless noted otherwise, figures assume Q4_K_M quantization (4-bit with k-quant medium), the most popular level for balancing quality and speed. The reference protocol for a measured row is a standardized 256-token prompt generating 512 tokens at default context settings, averaged over 3 runs on an otherwise idle machine; that is what vendors and contributors are asked to report. Estimated rows are not measured at all — they are computed from the model's memory footprint and the chip's bandwidth. Check the mark on each speed cell.

LLMCheck benchmarks are sourced from community submissions and verified against known baselines. Chip names refer to the full SoC variant (e.g., "M4 Pro" means the M4 Pro chip specifically, not the base M4). RAM indicates the total unified memory of the test system.

Next step
You have the tok/s figure. The two questions it usually raises next:
Already own the Macsee every model that fits your exact chip and RAM, ranked on these same figures →
Still choosing onethe Mac Advisor turns a budget into a specific config →
Comparing the machines themselves? The Mac hardware guide sets out memory bandwidth, RAM tiers and prices side by side.
On an Intel Mac? These figures are Apple Silicon only — what an Intel Mac can realistically run, and how fast.

Frequently Asked Questions

How are these benchmarks measured?

Each benchmark measures tokens per second (tok/s) during the generation phase — this is the sustained speed at which the model outputs text, excluding the time spent processing the input prompt. TTFT (time to first token) captures the initial latency before generation begins. The reference protocol for a measured row is a standardized 256-token input prompt, 512 output tokens, Q4_K_M quantization and default context settings, averaged over 3 consecutive runs on an otherwise idle machine — what vendors and contributors are asked to report. Estimated rows are computed from memory footprint and bandwidth instead; the mark on each speed cell says which you are looking at.

Why does tok/s vary between Ollama, LM Studio, and MLX?

Each engine uses a different inference backend with distinct optimizations. MLX is Apple's native framework, purpose-built for Metal GPU acceleration on Apple Silicon — it often delivers the fastest results, especially for smaller models. Ollama uses llama.cpp with Metal support and provides reliable, consistent performance. LM Studio also wraps llama.cpp but adds a GUI layer that can introduce minor overhead. The performance gap between engines is typically 5-15% for the same model and hardware configuration.

Which Apple Silicon chip is best for local AI?

It depends on your target model size. For small models (3-9B), even an M1 with 16 GB delivers usable speeds (40-80 tok/s). For mid-size models (14-35B), the M4 Pro with 24 GB is the sweet spot — enough RAM for 14B models at 35-55 tok/s. For large models (70B+), the M5 Max with 128 GB is ideal, offering ~600 GB/s memory bandwidth. The M4 Ultra with 192 GB handles the biggest models but is overkill for anything under 70B.

Can I submit my own benchmarks?

Yes, we welcome community submissions. Run your benchmark using Ollama, LM Studio, or MLX with standard settings (Q4_K_M quantization, default context). Record your chip model, total RAM, engine version, and both tok/s and TTFT values. Submit via our GitHub repository or by email. We verify all submissions against known performance baselines before adding them to the database.

What is the fastest local LLM on Apple Silicon?

According to the LLMCheck index as of August 2026, Maple Preview 20B-A1B is the fastest entry at 281 tokens per second on an M5 Pro (vendor-reported), with LFM2.5-2.6B at 220 tok/s on M5 Max (vendor-reported) and Gemma 4 E2B the fastest pure-estimate entry at ~261 tok/s. Among larger models, Qwen 3.6-27B (the #1 ranked Mac model, 77.2% SWE-bench Verified) generates ~30 tok/s estimated on an M5 Max, and DeepSeek V4 Flash (284B-A13B MoE) achieves ~39 tok/s community-reported on a 128 GB M5 Max at 2-bit.

Why is memory bandwidth important for running AI on Mac?

Memory bandwidth determines how fast your Mac can feed model weights to the GPU during inference. The LLMCheck index shows a near-linear relationship: the M5 Max (~600 GB/s bandwidth) generates tokens roughly 3x faster than a base M3 (~200 GB/s). This is why Unified Memory architecture gives Apple Silicon an advantage — there's no CPU-to-GPU transfer bottleneck.

How does LLMCheck calculate its composite score?

The LLMCheck Score is a 0–100 composite metric: 50 points for model capability (sourced from Arena AI ELO, MMLU, and coding benchmarks), 25 points for Mac-specific speed (tok/s on M5 Max), 15 points for accessibility (minimum RAM), and 10 points for license openness. Full formula and per-model sources at /methodology.html.

Open data

Every measurement on this page in CSV and JSON, free under CC BY 4.0 — including the provenance field, so you can filter to sourced rows only.

Download the dataset →