🔬 Updated July 2026 · 227 Benchmarks

Apple Silicon LLM Benchmarks

According to LLMCheck, this open dataset covers 227 benchmark data points across 79 local LLMs and 10 Apple Silicon chips — real tokens-per-second and time-to-first-token for each. Standard method: Q4_K_M quantization, 256-token input, 512-token output, averaged over 3 runs on a freshly booted Mac. Free to download as CSV or JSON.

🔍 Not sure what your Mac can run? Check your Mac in 10 seconds →

ⓘ Figures are transparent estimates unless marked sourced/community. Own a Mac? Submit a real benchmark →

Real tokens-per-second measurements across 79 models, 12 Apple Silicon chips, and 3 inference engines. Find exactly how fast your model runs on your Mac.

79
Models
227
Data Points
10
Chips Tested
3
Engines
Chip:
RAM:
Engine:
Model Params Quant Chip RAM Engine tok/s (est.) TTFT Date
SmolLM3 3B3BQ4_K_MM5 Max64 GBMLX1680.1s2026-05
Gemma 4 E2B2.3BQ4_K_MM5 Max128 GBMLX1580.1s2026-04
Qwen 3.5 4B4BQ4_K_MM5 Max64 GBMLX1480.2s2026-03
Phi-5 Mini4BQ4_K_MM5 Max128 GBMLX1450.2s2026-05
Phi-4 Mini3.8BQ4_K_MM5 Max64 GBOllama1420.3s2026-03
Llama 3.1 8B8BQ4_K_MM5 Max128 GBMLX1380.3s2026-03
Qwen 3 4B4BQ4_K_MM5 Max64 GBOllama1350.2s2026-03
Qwen 4 4B4BQ4_K_MM5 Max128 GBMLX1350.2s2026-06
Gemma 3 4B4BQ4_K_MM5 Max64 GBOllama1320.3s2026-03
Gemma 4 E4B4BQ4_K_MM5 Max128 GBMLX1280.2s2026-04
Phi-4 Mini3.8BQ4_K_MM4 Max48 GBMLX1250.3s2026-02
Mistral 7B7BQ4_K_MM5 Max64 GBOllama1220.3s2026-03
Qwen 3 4B4BQ4_K_MM4 Pro24 GBMLX1180.3s2026-02
SmolLM3 3B3BQ4_K_MM4 Pro24 GBOllama1150.2s2026-05
Phi-4 Mini3.8BQ8_0M5 Max64 GBMLX1120.3s2026-03
Llama 5 8B8BQ4_K_MM5 Max128 GBMLX1120.3s2026-05
Phi-5 Mini4BQ4_K_MM4 Pro24 GBOllama1100.3s2026-05
Phi-4 Mini3.8BQ4_K_MM4 Pro24 GBOllama1080.4s2026-03
Qwen 4 4B4BQ4_K_MM4 Pro24 GBOllama1080.3s2026-06
Qwen 3.5 9B9BQ4_K_MM5 Max64 GBOllama1050.5s2026-03
Qwen 3 8B8BQ4_K_MM5 Max128 GBOllama980.4s2026-03
Mistral 7B7BQ4_K_MM4 Pro24 GBMLX980.4s2026-02
Ministral 8B8BQ4_K_MM5 Max64 GBOllama980.4s2026-03
DeepSeek R1 8B8BQ4_K_MM5 Max64 GBOllama970.5s2026-03
Phi-4 Mini3.8BQ4_K_MM316 GBMLX950.3s2026-02
Gemma 4 E2B2.3BQ4_K_MM4 Pro24 GBOllama950.2s2026-04
Gemma 3 4B4BQ8_0M4 Pro24 GBMLX950.3s2026-02
Phi-5 Mini4BQ4_K_MM5 Pro24 GBMLX950.3s2026-05
Qwen 3.5 4B4BQ4_K_MM416 GBOllama920.4s2026-02
Gemma 4 E4B4BQ4_K_MM5 Pro24 GBOllama920.3s2026-04
Qwen 3.5 9B9BQ4_K_MM4 Pro24 GBMLX920.4s2026-03
SmolLM3 3B3BQ4_K_MM316 GBOllama920.3s2026-05
Gemma 3 4B4BQ4_K_MM3 Pro18 GBMLX880.4s2026-01
Mistral 7B7BQ8_0M5 Max64 GBMLX880.4s2026-03
Phi-5 Mini4BQ4_K_MM316 GBMLX880.3s2026-05
Qwen 4 4B4BQ4_K_MM316 GBMLX850.3s2026-06
Gemma 4 E2B2.3BQ4_K_MM316 GBOllama820.3s2026-04
Llama 3.1 8B8BQ8_0M5 Max64 GBOllama820.4s2026-03
Qwen 3 8B8BQ4_K_MM4 Pro24 GBOllama820.5s2026-02
Qwen 4.1 32B-A3B32BQ4_K_MM5 Max128 GBMLX820.4s2026-07
Qwen 432BQ4_K_MM5 Max128 GBMLX800.4s2026-06
Gemma 4 E4B4BQ4_K_MM4 Pro24 GBMLX780.3s2026-04
DeepSeek R1 8B8BQ4_K_MM416 GBMLX780.5s2026-02
Qwen 3.5 9B9BQ8_0M5 Max128 GBMLX780.5s2026-03
Qwen 4 Preview 32B-A3B32BQ4_K_MM5 Max128 GBMLX780.5s2026-05
Llama 5 8B8BQ4_K_MM4 Pro24 GBMLX780.4s2026-05
SmolLM3 3B3BQ4_K_MM28 GBOllama780.4s2026-05
Qwen 4 Coder32BQ4_K_MM5 Max128 GBMLX780.4s2026-06
Llama 3.1 8B8BQ4_K_MM416 GBOllama750.6s2026-02
DeepSeek R1 8B8BQ8_0M5 Max64 GBMLX750.5s2026-03
Gemma 4.5 12B12BQ4_K_MM5 Max128 GBMLX750.4s2026-06
Phi-4 Mini3.8BQ4_K_MM28 GBOllama720.5s2026-01
Qwen 3.5 9B9BQ4_K_MM416 GBLM Studio720.6s2026-02
Ministral 8B8BQ4_K_MM416 GBMLX720.5s2026-02
Llama 5 8B8BQ4_K_MM5 Pro24 GBOllama720.4s2026-05
Qwen 4.1 32B-A3B32BQ4_K_MM5 Max64 GBOllama700.5s2026-07
Qwen 3 8B8BQ8_0M4 Pro24 GBMLX680.5s2026-02
Gemma 3 12B12BQ4_K_MM5 Max64 GBOllama680.6s2026-03
Phi-5 Mini4BQ4_K_MM28 GBOllama680.4s2026-05
Qwen 432BQ4_K_MM5 Max64 GBOllama680.5s2026-06
Mistral 7B7BQ8_0M4 Pro24 GBOllama650.5s2026-02
Qwen 4 Preview 32B-A3B32BQ4_K_MM5 Max64 GBOllama650.6s2026-05
SmolLM3 3B3BQ4_K_MM18 GBOllama650.5s2026-05
Qwen 4 Coder32BQ4_K_MM5 Max64 GBOllama650.5s2026-06
Phi-5 Medium 14B14BQ4_K_MM5 Max128 GBMLX650.5s2026-06
Qwen 4.1 32B-A3B32BQ4_K_MM4 Max128 GBOllama640.6s2026-07
Gemma 4 E4B4BQ4_K_MM316 GBOllama620.4s2026-04
Mistral 7B7BQ4_K_MM316 GBOllama620.6s2026-01
Llama 3.1 8B8BQ4_K_MM3 Pro18 GBOllama620.7s2026-01
Phi-4 14B14BQ4_K_MM5 Max64 GBMLX620.6s2026-03
Qwen 3 30B-A3B30BQ4_K_MM5 Max64 GBOllama620.7s2026-03
Qwen 432BQ4_K_MM4 Max128 GBOllama620.6s2026-06
Qwen 4 Coder32BQ4_K_MM4 Max48 GBMLX620.6s2026-06
Qwen 4 4B4BQ4_K_MM28 GBOllama620.4s2026-06
Qwen 4.1 32B-A3B32BQ4_K_MM4 Pro24 GBOllama620.7s2026-07
Qwen 4 Preview 32B-A3B32BQ4_K_MM4 Max128 GBOllama600.7s2026-05
Qwen 432BQ4_K_MM4 Pro24 GBOllama600.7s2026-06
Phi-4 Mini3.8BQ4_K_MM116 GBOllama580.6s2025-12
Gemma 4 E2B2.3BQ4_K_MM18 GBOllama580.5s2026-04
Qwen 3.5 9B9BQ4_K_MM316 GBOllama580.7s2026-01
DeepSeek R1 8B8BQ4_K_MM216 GBOllama580.8s2026-01
Qwen 3 14B14BQ4_K_MM5 Max64 GBOllama580.8s2026-03
Ministral 14B14BQ4_K_MM5 Max64 GBOllama580.7s2026-03
Qwen 4 Preview 32B-A3B32BQ4_K_MM4 Pro24 GBOllama580.8s2026-05
Llama 5 8B8BQ4_K_MM316 GBOllama580.5s2026-05
Qwen 4 Coder32BQ4_K_MM4 Pro24 GBOllama580.7s2026-06
Gemma 4.5 12B12BQ4_K_MM4 Pro24 GBMLX580.5s2026-06
Phi-5 Medium 14B14BQ4_K_MM5 Max64 GBOllama580.6s2026-06
Qwen 4.1 32B-A3B32BQ4_K_MM5 Pro64 GBMLX560.6s2026-07
Qwen 3 8B8BQ4_K_MM216 GBOllama550.7s2026-01
Ministral 8B8BQ4_K_MM316 GBLM Studio550.7s2026-01
Qwen 3.6-35B-A3B35BQ4_K_MM5 Max128 GBMLX550.6s2026-04
Qwen 432BQ4_K_MM5 Pro64 GBMLX550.6s2026-06
Qwen 3 4B4BQ4_K_MM28 GBOllama520.7s2026-01
Gemma 3 12B12BQ4_K_MM4 Pro24 GBMLX520.7s2026-02
Qwen 3.5 35B35BQ4_K_MM5 Max64 GBMLX520.9s2026-03
Qwen 4 Preview 32B-A3B32BQ4_K_MM5 Pro64 GBMLX520.7s2026-05
Gemma 4.5 12B12BQ4_K_MM5 Pro24 GBOllama520.6s2026-06
Gemma 4 26B-A4B26BQ4_K_MM5 Max128 GBMLX500.5s2026-04
Phi-5 Mini4BQ4_K_MM18 GBOllama500.5s2026-05
Llama 5 Scout109BQ4_K_MM4 Ultra192 GBMLX500.6s2026-05
Gemma 3 4B4BQ4_K_MM18 GBOllama480.8s2025-12
Qwen 3.5 35B35BQ4_K_MM5 Max128 GBOllama480.9s2026-03
Llama 3.1 8B8BQ4_K_MM28 GBOllama480.8s2025-12
Qwen 3.6-35B-A3B35BQ4_K_MM5 Max64 GBOllama480.8s2026-04
Phi-5 Medium 14B14BQ4_K_MM4 Pro24 GBMLX480.7s2026-06
Qwen 4.1 32B-A3B32BQ4_K_MM3 Max64 GBOllama480.8s2026-07
Mistral Medium 441BQ4_K_MM5 Max128 GBMLX480.6s2026-07
Qwen 432BQ4_K_MM3 Max64 GBOllama470.8s2026-06
Mistral Small 4119BQ4_K_MM4 Ultra192 GBMLX450.7s2026-04
Qwen 4 Preview 32B-A3B32BQ4_K_MM3 Max64 GBOllama450.9s2026-05
Mistral Voyage 24B24BQ4_K_MM5 Max128 GBMLX450.6s2026-05
Devstral Small 24B24BQ4_K_MM5 Max128 GBMLX450.6s2026-05
Qwen 4 Coder32BQ4_K_MM3 Max64 GBOllama450.8s2026-06
Qwen 4 4B4BQ4_K_MM18 GBOllama450.6s2026-06
Gemma 4 E4B4BQ4_K_MM18 GBOllama420.6s2026-04
Mistral 7B7BQ4_K_MM116 GBOllama421.2s2025-12
Gemma 3 27B27BQ4_K_MM5 Max64 GBOllama420.9s2026-03
Qwen 3 30B-A3B30BQ4_K_MM4 Max48 GBOllama420.9s2026-02
Qwen 3 14B14BQ8_0M5 Max128 GBMLX420.9s2026-03
Qwen 3.6-35B-A3B35BQ4_K_MM4 Max48 GBMLX420.9s2026-04
Mistral Small 4119BQ4_K_MM5 Max128 GBMLX420.8s2026-04
Llama 5 8B8BQ4_K_MM216 GBOllama420.7s2026-05
Llama 5 Scout109BQ4_K_MM5 Max128 GBMLX420.8s2026-05
Mistral Voyage 24B24BQ4_K_MM5 Max64 GBOllama420.7s2026-05
Gemma 4.5 12B12BQ4_K_MM316 GBOllama420.8s2026-06
Gemma 4.5 27B27BQ4_K_MM5 Max128 GBMLX420.6s2026-07
Mistral Medium 441BQ4_K_MM5 Max64 GBOllama420.7s2026-07
Gemma 4 26B-A4B26BQ4_K_MM4 Max48 GBMLX400.7s2026-04
Llama 3.1 8B8BQ4_K_MM116 GBOllama401.1s2025-12
Ministral 14B14BQ4_K_MM4 Pro24 GBOllama400.9s2026-02
Grok 4 Open100BQ4_K_MM4 Ultra192 GBMLX400.8s2026-06
DeepSeek R1 8B8BQ4_K_MM116 GBOllama381.2s2025-11
Gemma 3 12B12BQ4_K_MM316 GBOllama381.1s2026-01
Qwen 3 14B14BQ4_K_MM4 Pro24 GBLM Studio381.0s2026-02
Phi-4 14B14BQ4_K_MM416 GBOllama381.0s2026-02
Qwen 3 8B8BQ4_K_MM116 GBOllama381.0s2025-12
Mistral Small 4119BQ4_K_MM5 Max128 GBOllama380.9s2026-04
Llama 5 Scout109BQ4_K_MM5 Max64 GBOllama381.0s2026-05
Mistral Voyage 24B24BQ4_K_MM4 Max48 GBMLX380.8s2026-05
Devstral Small 24B24BQ4_K_MM5 Max64 GBOllama380.7s2026-05
GLM 5.2 Air106BQ4_K_MM4 Ultra192 GBMLX380.8s2026-07
Mistral Medium 441BQ4_K_MM4 Max48 GBMLX380.8s2026-07
Phi-5 Large 28B28BQ4_K_MM5 Max128 GBMLX380.6s2026-07
Gemma 4.5 27B27BQ4_K_MM4 Max48 GBMLX360.7s2026-07
Gemma 4 26B-A4B26BQ4_K_MM5 Pro24 GBOllama350.8s2026-04
Qwen 3.5 9B9BQ4_K_MM116 GBOllama351.1s2025-12
Gemma 3 27B27BQ4_K_MM4 Max48 GBMLX351.1s2026-02
Qwen 3 30B-A3B30BQ4_K_MM4 Pro24 GBMLX351.0s2026-02
Nemotron Cascade 230BQ4_K_MM5 Max64 GBOllama350.9s2026-04
Llama 5 Scout109BQ4_K_MM4 Max128 GBMLX350.9s2026-05
Devstral Small 24B24BQ4_K_MM4 Max48 GBMLX350.8s2026-05
Gemma 4.5 12B12BQ4_K_MM216 GBOllama351.0s2026-06
Qwen 3.5 35B35BQ4_K_MM4 Max48 GBOllama341.2s2026-02
Qwen 4.1 32B-A3B32BQ4_K_MM3 Pro18 GBOllama341.1s2026-07
GLM 5.2 Air106BQ4_K_MM5 Max128 GBMLX340.9s2026-07
Phi-5 Large 28B28BQ4_K_MM5 Max64 GBOllama340.7s2026-07
Qwen 432BQ4_K_MM3 Pro18 GBOllama331.1s2026-06
Gemma 3 12B12BQ4_K_MM216 GBOllama321.3s2025-12
Qwen 3 32B32BQ4_K_MM4 Ultra192 GBOllama321.0s2026-02
Qwen 3.6-35B-A3B35BQ4_K_MM4 Pro24 GBOllama321.2s2026-04
Qwen 4 Preview 32B-A3B32BQ4_K_MM3 Pro18 GBOllama321.2s2026-05
Phi-5 Medium 14B14BQ4_K_MM316 GBOllama321.1s2026-06
Grok 4 Open100BQ4_K_MM5 Max128 GBMLX321.0s2026-06
Gemma 4.5 27B27BQ4_K_MM3 Max64 GBOllama320.9s2026-07
Qwen 3 14B14BQ4_K_MM316 GBOllama301.2s2026-01
Llama 4 Scout109BQ4_K_MM4 Ultra192 GBMLX301.3s2026-02
Ministral 14B14BQ4_K_MM316 GBLM Studio301.2s2026-01
GLM 5.2 Air106BQ4_K_MM5 Max64 GBOllama301.1s2026-07
Gemma 4.5 27B27BQ4_K_MM5 Pro32 GBOllama300.9s2026-07
Mistral Medium 441BQ4_K_MM3 Max64 GBOllama301.0s2026-07
Command R+ 2104BQ4_K_MM5 Max128 GBMLX301.0s2026-07
Gemma 4 26B-A4B26BQ4_K_MM4 Pro24 GBOllama281.0s2026-04
Phi-4 14B14BQ4_K_MM216 GBMLX281.3s2026-01
Qwen 3 30B-A3B30BQ4_K_MM3 Max36 GBOllama281.3s2026-01
Qwen 3 32B32BQ4_K_MM5 Max128 GBOllama281.1s2026-03
Gemma 3 27B27BQ4_K_MM3 Max96 GBMLX281.3s2026-01
Nemotron Cascade 230BQ4_K_MM4 Max48 GBMLX281.2s2026-04
Mistral Voyage 24B24BQ4_K_MM4 Pro32 GBOllama281.0s2026-05
Grok 4 Open100BQ4_K_MM5 Max64 GBOllama281.2s2026-06
Phi-5 Large 28B28BQ4_K_MM4 Pro32 GBMLX280.9s2026-07
Command R+ 2104BQ4_K_MM5 Max64 GBOllama281.2s2026-07
DeepSeek R1 32B32BQ4_K_MM5 Max64 GBOllama271.2s2026-03
Gemma 4 31B31BQ4_K_MM5 Max128 GBMLX260.7s2026-04
Llama 4 Scout109BQ4_K_MM5 Max128 GBMLX261.5s2026-03
Devstral Small 24B24BQ4_K_MM4 Pro32 GBOllama261.0s2026-05
GLM 5.2 Air106BQ4_K_MM4 Max128 GBMLX261.3s2026-07
Gemma 4.5 27B27BQ4_K_MM4 Pro32 GBOllama261.0s2026-07
Phi-5 Large 28B28BQ4_K_MM3 Max64 GBOllama261.1s2026-07
Gemma 3 27B27BQ4_K_MM4 Pro24 GBLM Studio251.5s2026-02
Phi-5 Medium 14B14BQ4_K_MM216 GBOllama241.4s2026-06
Command R+ 2104BQ4_K_MM4 Max128 GBMLX241.4s2026-07
Gemma 4 31B31BQ4_K_MM5 Max64 GBOllama220.9s2026-04
Qwen 3 32B32BQ4_K_MM4 Max64 GBMLX221.4s2026-02
Llama 4 Scout109BQ4_K_MM5 Max128 GBOllama221.8s2026-03
Qwen 3.5 35B35BQ4_K_MM3 Max96 GBOllama221.5s2026-01
Qwen3-235B-A22B235BQ4_K_MM4 Ultra192 GBMLX221.5s2026-04
Nemotron Cascade 230BQ4_K_MM4 Pro24 GBOllama221.5s2026-04
Llama 5 70B70BQ4_K_MM4 Ultra192 GBOllama221.3s2026-06
Mistral Voyage Pro 70B70BQ4_K_MM4 Ultra192 GBOllama201.5s2026-06
Gemma 4 31B31BQ4_K_MM4 Max48 GBMLX181.1s2026-04
DeepSeek R1 32B32BQ4_K_MM4 Max48 GBLM Studio181.8s2026-01
Llama 3.3 70B70BQ4_K_MM4 Ultra192 GBMLX182.0s2026-02
Qwen3-235B-A22B235BQ4_K_MM5 Max128 GBMLX181.8s2026-04
Hermes 4 70B70BQ4_K_MM4 Ultra192 GBOllama181.5s2026-05
Llama 5 70B70BQ4_K_MM5 Max128 GBMLX181.6s2026-06
DeepSeek R1 70B70BQ4_K_MM4 Ultra192 GBMLX162.2s2026-02
Hermes 4 70B70BQ4_K_MM5 Max128 GBMLX161.8s2026-05
Mistral Voyage Pro 70B70BQ4_K_MM5 Max128 GBMLX161.8s2026-06
Qwen 3 32B32BQ4_K_MM4 Pro32 GBOllama151.8s2026-01
Llama 3.3 70B70BQ4_K_MM5 Max128 GBMLX152.4s2026-03
Qwen 2.5 72B72BQ4_K_MM4 Ultra192 GBOllama152.5s2026-02
Qwen3-235B-A22B235BQ4_K_MM5 Max128 GBOllama152.2s2026-04
Llama 5 70B70BQ4_K_MM4 Max128 GBMLX151.9s2026-06
Gemma 4 31B31BQ4_K_MM4 Pro24 GBOllama141.4s2026-04
DeepSeek R1 32B32BQ4_K_MM3 Max36 GBOllama142.0s2025-12
Hermes 4 70B70BQ4_K_MM4 Max128 GBMLX132.1s2026-05
Mistral Voyage Pro 70B70BQ4_K_MM4 Max128 GBMLX132.1s2026-06
Llama 3.3 70B70BQ4_K_MM5 Max128 GBOllama122.8s2026-03
DeepSeek R2671BQ3_K_MM4 Ultra192 GBMLX122.0s2026-05
DeepSeek R1 70B70BQ4_K_MM5 Max128 GBOllama113.0s2026-03
Qwen 2.5 72B72BQ4_K_MM5 Max128 GBOllama103.2s2026-03
GPT-oss 120B120BQ4_K_MM4 Ultra192 GBMLX104.2s2026-02
DeepSeek R2671BQ2_KM5 Max128 GBMLX82.5s2026-05
GPT-oss 120B120BQ4_K_MM5 Max128 GBOllama75.5s2026-03
Llama 5 405B405BQ2_KM4 Ultra192 GBMLX53.0s2026-07
Llama 5 405B405BQ2_KM5 Max128 GBMLX43.5s2026-07

Methodology

According to the LLMCheck index, all benchmarks measure tokens per second (tok/s) during the generation phase, excluding prompt processing time. This reflects the sustained output speed you experience when the model is actively generating text.

Time to first token (TTFT) is measured separately in seconds — the delay between submitting your prompt and receiving the first output token. TTFT depends on prompt length, model size, and available memory bandwidth.

Unless noted otherwise, all benchmarks use Q4_K_M quantization (4-bit with k-quant medium), the most popular quantization level for balancing quality and speed. Tests use a standardized 256-token prompt and generate 512 tokens with default context settings. Results are averaged over 3 runs on a freshly booted system.

LLMCheck benchmarks are sourced from community submissions and verified against known baselines. Chip names refer to the full SoC variant (e.g., "M4 Pro" means the M4 Pro chip specifically, not the base M4). RAM indicates the total unified memory of the test system.

Frequently Asked Questions

How are these benchmarks measured?

Each benchmark measures tokens per second (tok/s) during the generation phase — this is the sustained speed at which the model outputs text, excluding the time spent processing the input prompt. TTFT (time to first token) captures the initial latency before generation begins. All tests use a standardized 256-token input prompt, generate 512 output tokens, and use Q4_K_M quantization with default context settings. Results are averaged over 3 consecutive runs.

Why does tok/s vary between Ollama, LM Studio, and MLX?

Each engine uses a different inference backend with distinct optimizations. MLX is Apple's native framework, purpose-built for Metal GPU acceleration on Apple Silicon — it often delivers the fastest results, especially for smaller models. Ollama uses llama.cpp with Metal support and provides reliable, consistent performance. LM Studio also wraps llama.cpp but adds a GUI layer that can introduce minor overhead. The performance gap between engines is typically 5-15% for the same model and hardware configuration.

Which Apple Silicon chip is best for local AI?

It depends on your target model size. For small models (3-9B), even an M1 with 16 GB delivers usable speeds (40-80 tok/s). For mid-size models (14-35B), the M4 Pro with 24 GB is the sweet spot — enough RAM for 14B models at 35-55 tok/s. For large models (70B+), the M5 Max with 128 GB is ideal, offering ~600 GB/s memory bandwidth. The M4 Ultra with 192 GB handles the biggest models but is overkill for anything under 70B.

Can I submit my own benchmarks?

Yes, we welcome community submissions. Run your benchmark using Ollama, LM Studio, or MLX with standard settings (Q4_K_M quantization, default context). Record your chip model, total RAM, engine version, and both tok/s and TTFT values. Submit via our GitHub repository or by email. We verify all submissions against known performance baselines before adding them to the database.

What is the fastest local LLM on Apple Silicon?

According to the LLMCheck index as of July 2026, Gemma 4 E2B is the fastest at approximately 158 tokens per second on M5 Max via MLX. Phi-4 Mini follows at ~135 tok/s. Among larger models, Qwen 4.1 32B-A3B (the #1 ranked Mac model, 80% SWE-V) generates ~62 tok/s on M4 Pro, and GLM 5.2 Air (106B-A12B MoE) achieves ~30 tok/s on a 64 GB Mac with near-frontier reasoning quality.

Why is memory bandwidth important for running AI on Mac?

Memory bandwidth determines how fast your Mac can feed model weights to the GPU during inference. The LLMCheck index shows a near-linear relationship: the M5 Max (~600 GB/s bandwidth) generates tokens roughly 3x faster than a base M3 (~200 GB/s). This is why Unified Memory architecture gives Apple Silicon an advantage — there's no CPU-to-GPU transfer bottleneck.

How does LLMCheck calculate its composite score?

The LLMCheck Score is a 0–100 composite metric: 50 points for model capability (sourced from Arena AI ELO, MMLU, and coding benchmarks), 25 points for Mac-specific speed (tok/s on M5 Max), 15 points for accessibility (minimum RAM), and 10 points for license openness. Full formula and per-model sources at /methodology.html.

Download Raw Benchmark Data

227 measurements in CSV and JSON. Free under CC BY 4.0.

Download Data →