Updated August 2026 · 202 Benchmarks

Apple Silicon LLM Benchmarks

According to LLMCheck, this open dataset covers 244 benchmark data points across 66 local LLMs and 15 Apple Silicon chips — tokens-per-second for each, and time-to-first-token where one was published. Every figure is labelled with its provenance: 11 published figures (vendor or press review), 13 linked individual runs, and 220 derived from the published memory-bandwidth model at each row's stated quantization. Free to download as CSV or JSON.

🔍 Not sure what your Mac can run? Check your Mac in 10 seconds →

ⓘ Figures are transparent estimates unless marked sourced/community. Own a Mac? Submit a real benchmark →

Tokens-per-second figures — estimated, sourced, and community-submitted, each row labeled — across 56 models, 16 Apple Silicon chips, and 3 inference engines. Find exactly how fast your model runs on your Mac.

56
Models
202
Data Points
14
Chips Tested
3
Engines
Chip:
RAM:
Engine:
Model↕ Params↕ Quant↕ Chip↕ RAM↕ Engine↕ tok/s (est.)↓ TTFT↕ Date↕
Phi-4 Mini3.8BQ4_K_MM5 Max64 GBOllamaestimated168—2026-03
Phi-4 Mini3.8BQ4_K_MM4 Pro24 GBOllamaestimated86—2026-03
Phi-4 Mini3.8BQ4_K_MM316 GBMLXestimated34—2026-02
Phi-4 Mini3.8BQ4_K_MM28 GBOllamaestimated34—2026-01
Phi-4 Mini3.8BQ4_K_MM116 GBOllamaestimated24—2025-12
Qwen 3.5 4B4BQ4_K_MM5 Max64 GBMLXestimated161—2026-03
Qwen 3.5 4B4BQ4_K_MM416 GBOllamaestimated39—2026-02
Qwen 3 4B4BQ4_K_MM28 GBOllamaestimated33—2026-01
Qwen 3 4B4BQ4_K_MM4 Pro24 GBMLXestimated82—2026-02
Gemma 4 E2B2.3BQ4_K_MM5 Max128 GBMLXestimated261—2026-04
Gemma 4 E2B2.3BQ4_K_MM4 Pro24 GBOllamaestimated147—2026-04
Gemma 4 E2B2.3BQ4_K_MM316 GBOllamaestimated62—2026-04
Gemma 4 E2B2.3BQ4_K_MM18 GBOllamaestimated44—2026-04
Gemma 4 E4B4BQ4_K_MM5 Max128 GBMLXestimated161—2026-04
Gemma 4 E4B4BQ4_K_MM5 Pro24 GBOllamaestimated91—2026-04
Gemma 4 E4B4BQ4_K_MM4 Pro24 GBMLXestimated82—2026-04
Gemma 4 E4B4BQ4_K_MM316 GBOllamaestimated33—2026-04
Gemma 4 E4B4BQ4_K_MM18 GBOllamaestimated23—2026-04
Gemma 4 26B-A4B26BQ4_K_MM5 Max128 GBMLXestimated106—2026-04
Gemma 4 26B-A4B26BQ4_K_MM5 Pro24 GBOllamaestimated72—2026-04
Gemma 4 26B-A4B26BQ4_K_MM4 Max48 GBMLXestimated100—2026-04
Gemma 4 26B-A4B26BQ4_K_MM4 Pro24 GBOllamaestimated67—2026-04
Gemma 4 31B31BQ4_K_MM5 Max128 GBMLXestimated26—2026-04
Gemma 4 31B31BQ4_K_MM4 Max48 GBMLXestimated23—2026-04
Gemma 4 31B31BQ4_K_MM4 Pro24 GBOllamaestimated12—2026-04
Gemma 3 4B4BQ4_K_MM5 Max64 GBOllamaestimated161—2026-03
Gemma 3 4B4BQ4_K_MM18 GBOllamaestimated23—2025-12
Gemma 3 4B4BQ4_K_MM3 Pro18 GBMLXestimated48—2026-01
Qwen 3.5 9B9BQ4_K_MM5 Max64 GBOllamaestimated82—2026-03
Qwen 3.5 9B9BQ4_K_MM416 GBLM Studioestimated18—2026-02
Qwen 3.5 9B9BQ4_K_MM316 GBOllamaestimated15—2026-01
Qwen 3.5 9B9BQ4_K_MM116 GBOllamaestimated10—2025-12
Qwen 3 8B8BQ4_K_MM5 Max128 GBOllamaestimated91—2026-03
Qwen 3 8B8BQ8_0M4 Pro24 GBMLXestimated24—2026-02
Qwen 3 8B8BQ4_K_MM216 GBOllamaestimated17—2026-01
DeepSeek R1 8B8BQ4_K_MM5 Max64 GBOllamaestimated91—2026-03
DeepSeek R1 8B8BQ4_K_MM416 GBMLXestimated20—2026-02
DeepSeek R1 8B8BQ4_K_MM216 GBOllamaestimated17—2026-01
DeepSeek R1 8B8BQ4_K_MM116 GBOllamaestimated11—2025-11
Mistral 7B7BQ4_K_MM5 Max64 GBOllamaestimated102—2026-03
Mistral 7B7BQ4_K_MM4 Pro24 GBMLXestimated50—2026-02
Mistral 7B7BQ4_K_MM316 GBOllamaestimated19—2026-01
Mistral 7B7BQ4_K_MM116 GBOllamaestimated13—2025-12
Llama 3.1 8B8BQ4_K_MM5 Max128 GBMLXestimated91—2026-03
Llama 3.1 8B8BQ4_K_MM416 GBOllamaestimated20—2026-02
Llama 3.1 8B8BQ4_K_MM3 Pro18 GBOllamaestimated25—2026-01
Llama 3.1 8B8BQ4_K_MM116 GBOllamaestimated11—2025-12
Ministral 8B8BQ4_K_MM5 Max64 GBOllamaestimated91—2026-03
Ministral 8B8BQ4_K_MM316 GBLM Studioestimated17—2026-01
Gemma 3 12B12BQ4_K_MM5 Max64 GBOllamaestimated64—2026-03
Gemma 3 12B12BQ4_K_MM4 Pro24 GBMLXestimated30—2026-02
Gemma 3 12B12BQ4_K_MM316 GBOllamaestimated11—2026-01
Gemma 3 12B12BQ4_K_MM216 GBOllamaestimated11—2025-12
Qwen 3 14B14BQ4_K_MM5 Max64 GBOllamaestimated55—2026-03
Qwen 3 14B14BQ4_K_MM4 Pro24 GBLM Studioestimated26—2026-02
Qwen 3 14B14BQ4_K_MM316 GBOllamaestimated10—2026-01
Phi-4 14B14BQ4_K_MM5 Max64 GBMLXestimated55—2026-03
Phi-4 14B14BQ4_K_MM416 GBOllamaestimated12—2026-02
Phi-4 14B14BQ4_K_MM216 GBMLXestimated10—2026-01
Ministral 3 14B14BQ4_K_MM5 Max64 GBOllamaestimated55—2026-03
Ministral 3 14B14BQ4_K_MM4 Pro24 GBOllamaestimated26—2026-02
Gemma 3 27B27BQ4_K_MM5 Max64 GBOllamaestimated30—2026-03
Gemma 3 27B27BQ4_K_MM4 Pro24 GBLM Studioestimated14—2026-02
Gemma 3 27B27BQ4_K_MM4 Max48 GBMLXestimated27—2026-02
Qwen 3 30B-A3B30BQ4_K_MM5 Max64 GBOllamaestimated116—2026-03
Qwen 3 30B-A3B30BQ4_K_MM4 Pro24 GBMLXestimated76—2026-02
Qwen 3 30B-A3B30BQ4_K_MM3 Max36 GBOllamaestimated94—2026-01
Qwen 3.5 35B-A3B35BQ4_K_MM5 Max64 GBMLXestimated117—2026-03
Qwen 3.5 35B-A3B35BQ4_K_MM4 Max48 GBOllamaestimated111—2026-02
Qwen 3 32B32BQ4_K_MM5 Max128 GBOllamaestimated25—2026-03
Qwen 3 32B32BQ4_K_MM4 Max64 GBMLXestimated23—2026-02
Qwen 3 32B32BQ4_K_MM4 Pro32 GBOllamaestimated12—2026-01
DeepSeek R1 32B32BQ4_K_MM5 Max64 GBOllamaestimated25—2026-03
DeepSeek R1 32B32BQ4_K_MM4 Max48 GBLM Studioestimated23—2026-01
DeepSeek R1 32B32BQ4_K_MM3 Max36 GBOllamaestimated17—2025-12
Llama 3.3 70B70BQ4_K_MM5 Max128 GBMLXestimated12—2026-03
DeepSeek R1 70B70BQ4_K_MM5 Max128 GBOllamaestimated12—2026-03
Qwen 2.5 72B72BQ4_K_MM5 Max128 GBOllamaestimated12—2026-03
GPT-oss 120B117B-A5.1BMXFP4M5 Max128 GBOllamaestimated58—2026-03
Llama 4 Scout109BQ4_K_MM5 Max128 GBMLXestimated19—2026-03
Mistral 7B7BQ8_0M5 Max64 GBMLXestimated59—2026-03
Mistral 7B7BQ8_0M4 Pro24 GBOllamaestimated28—2026-02
Qwen 3.5 9B9BQ8_0M5 Max128 GBMLXestimated47—2026-03
Llama 3.1 8B8BQ8_0M5 Max64 GBOllamaestimated52—2026-03
Qwen 3 14B14BQ8_0M5 Max128 GBMLXestimated31—2026-03
Phi-4 Mini3.8BQ8_0M5 Max64 GBMLXestimated102—2026-03
Phi-4 Mini3.8BQ4_K_MM4 Max48 GBMLXestimated153—2026-02
Gemma 3 4B4BQ8_0M4 Pro24 GBMLXestimated47—2026-02
Qwen 3.5 35B-A3B35BQ4_K_MM3 Max96 GBOllamaestimated96—2026-01
DeepSeek R1 8B8BQ8_0M5 Max64 GBMLXestimated52—2026-03
Qwen 3 8B8BQ4_K_MM4 Pro24 GBOllamaestimated44—2026-02
Qwen 3 8B8BQ4_K_MM116 GBOllamaestimated11—2025-12
Ministral 8B8BQ4_K_MM416 GBMLXestimated20—2026-02
Ministral 3 14B14BQ4_K_MM316 GBLM Studioestimated10—2026-01
Gemma 3 27B27BQ4_K_MM3 Max96 GBMLXestimated20—2026-01
Llama 3.1 8B8BQ4_K_MM28 GBOllamaestimated17—2025-12
Qwen 3.5 9B9BQ4_K_MM4 Pro24 GBMLXestimated39—2026-03
Qwen 3 4B4BQ4_K_MM5 Max64 GBOllamaestimated161—2026-03
Qwen 3.6-35B-A3B35BQ4_K_MM4 Max48 GBMLXestimated120—2026-04
Qwen 3.6-35B-A3B35BQ4_K_MM4 Pro24 GBOllamaestimated86—2026-04
Mistral Small 4119BQ4_K_MM5 Max128 GBMLXestimated46—2026-04
Qwen3-235B-A22B235BQ4_K_MM5 Max128 GBMLXestimated15—2026-04
Nemotron-Cascade 230BQ4_K_MM5 Max64 GBOllamaestimated117—2026-04
Nemotron-Cascade 230BQ4_K_MM4 Max48 GBMLXestimated111—2026-04
Nemotron-Cascade 230BQ4_K_MM4 Pro24 GBOllamaestimated77—2026-04
Mistral Small 3.2 24B24BQ4_K_MM5 Max128 GBMLXestimated33—2026-05
Mistral Small 3.2 24B24BQ4_K_MM4 Max48 GBMLXestimated30—2026-05
Mistral Small 3.2 24B24BQ4_K_MM4 Pro32 GBOllamaestimated15—2026-05
Hermes 4 70B70BQ4_K_MM5 Max128 GBMLXestimated12—2026-05
Hermes 4 70B70BQ4_K_MM4 Max128 GBMLXestimated11—2026-05
SmolLM3 3B3BQ4_K_MM5 Max64 GBMLXestimated199—2026-05
SmolLM3 3B3BQ4_K_MM4 Pro24 GBOllamaestimated106—2026-05
SmolLM3 3B3BQ4_K_MM316 GBOllamaestimated43—2026-05
SmolLM3 3B3BQ4_K_MM28 GBOllamaestimated43—2026-05
SmolLM3 3B3BQ4_K_MM18 GBOllamaestimated30—2026-05
Devstral Small 24B24BQ4_K_MM5 Max128 GBMLXestimated33—2026-05
Devstral Small 24B24BQ4_K_MM4 Max48 GBMLXestimated30—2026-05
Devstral Small 24B24BQ4_K_MM4 Pro32 GBOllamaestimated15—2026-05
GLM-4.5-Air106BQ4_K_MM5 Max128 GBMLXestimated27—2026-07
GLM-4.5-Air106BQ4_K_MM4 Max128 GBMLXestimated24—2026-07
DeepSeek V4 Flash284B-A13BQ2_KM5 Max128 GBMLXcommunity390.4s2026-08
DeepSeek V4 Flash284B-A13BQ2_KM3 Max128 GBMLXcommunity270.6s2026-08
GLM 5.2753B-A40BIQ1_SM3 Ultra256 GBLM Studiocommunity221.5s2026-07
GLM 5.2753B-A40BQ4_K_MM3 Ultra512 GBMLXcommunity152.0s2026-07
Muse Glimmer 30B30BQ4_K_MM4 Pro24 GBLM Studiocommunity100.8s2026-08
Muse Glimmer 30B30BQ4_K_MM5 Max64 GBMLXsourced270.4s2026-08
Muse Glimmer 30B30BQ4_K_MM4 Max48 GBMLXsourced240.5s2026-08
LFM2.5-2.6B2.6BQ4_K_MM5 Max64 GBMLXsourced2200.1s2026-08
Maple Preview 20B-A1B20B-A1BTernaryM5 Pro24 GBMLXsourced2810.1s2026-08
Maple Preview 20B-A1B20B-A1BTernaryM416 GBMLXsourced2000.1s2026-08
Qwen 3.6-27B27BQ4_K_MM5 Max64 GBMLXestimated30—2026-08
Qwen 3.6-27B27BQ4_K_MM4 Max64 GBMLXestimated27—2026-08
Qwen 3.6-27B27BQ4_K_MM3 Max64 GBOllamaestimated20—2026-08
Qwen 3.6-27B27BQ4_K_MM5 Pro24 GBMLXestimated15—2026-08
Qwen 3.6-27B27BQ4_K_MM4 Pro24 GBOllamaestimated14—2026-08
KAT-Coder-V2.535B-A3BQ4_K_MM5 Max64 GBMLXestimated117—2026-08
KAT-Coder-V2.535B-A3BQ4_K_MM4 Max48 GBMLXestimated111—2026-08
KAT-Coder-V2.535B-A3BQ4_K_MM4 Pro24 GBOllamaestimated77—2026-08
Nemotron 3.5 Lightning30B-A3BQ4_K_MM5 Max64 GBMLXestimated103—2026-08
Nemotron 3.5 Lightning30B-A3BQ4_K_MM4 Max48 GBMLXestimated97—2026-08
Nemotron 3.5 Lightning30B-A3BQ4_K_MM4 Pro24 GBOllamaestimated64—2026-08
Laguna S 2.1118B-A8BQ4_K_MM5 Max128 GBMLXestimated38—2026-08
Laguna S 2.1118B-A8BQ4_K_MM4 Max128 GBMLXestimated35—2026-08
Laguna XS 2.133B-A3BQ4_K_MM5 Max64 GBMLXestimated117—2026-08
Laguna XS 2.133B-A3BQ4_K_MM4 Pro24 GBOllamaestimated77—2026-08
Inkling-Small276B-A12BQ2_KM5 Max128 GBMLXestimated42—2026-08
Bonsai 27B27B1-bitM5 Max64 GBMLXestimated30—2026-08
Bonsai 27B27B1-bitM416 GBMLXestimated6—2026-08
Bonsai 27B27B1-bitM28 GBMLXestimated5—2026-08
Nanbeige4.2-3B3BQ4_K_MM5 Max64 GBMLXestimated199—2026-08
Nanbeige4.2-3B3BQ4_K_MM416 GBOllamaestimated51—2026-08
Nanbeige4.2-3B3BQ4_K_MM28 GBOllamaestimated43—2026-08
Apertus 1.5 8B8BQ4_K_MM5 Max64 GBMLXestimated91—2026-08
Apertus 1.5 8B8BQ4_K_MM416 GBOllamaestimated20—2026-08
Apertus 1.5 70B70BQ4_K_MM5 Max128 GBMLXestimated12—2026-08
GLM-4.7-Flash31BQ4_K_MM5 Max64 GBMLXestimated26—2026-08
GLM-4.7-Flash31BQ4_K_MM4 Max48 GBMLXestimated23—2026-08
GLM-4.7-Flash31BQ4_K_MM4 Pro24 GBOllamaestimated12—2026-08
Llama 3.3 70B70BQ4_K_MM1 Ultra128 GBOllamaestimated15—2026-08
DeepSeek R1 70B70BQ4_K_MM1 Ultra128 GBOllamaestimated15—2026-08
Qwen 3.6-35B-A3B35B-A3BQ4_K_MM1 Ultra64 GBMLXestimated138—2026-08
Qwen 3.6-27B27BQ4_K_MM1 Ultra64 GBMLXestimated38—2026-08
Gemma 4 26B-A4B26B-A4BQ4_K_MM1 Ultra64 GBOllamaestimated119—2026-08
GPT-oss 120B117B-A5.1BMXFP4M1 Ultra128 GBOllamaestimated72—2026-08
Llama 3.3 70B70BQ4_K_MM2 Max96 GBOllamaestimated8—2026-08
Qwen 3.6-35B-A3B35B-A3BQ4_K_MM2 Max64 GBMLXestimated105—2026-08
Qwen 3.6-27B27BQ4_K_MM2 Max32 GBMLXestimated20—2026-08
Gemma 3 27B27BQ4_K_MM2 Max32 GBOllamaestimated20—2026-08
Llama 3.1 8B8BQ4_K_MM2 Max32 GBOllamaestimated62—2026-08
Mistral Small 3.2 24B24BQ4_K_MM2 Max32 GBOllamaestimated22—2026-08
Llama 3.3 70B70BQ4_K_MM3 Ultra256 GBOllamaestimated16—2026-08
DeepSeek V4 Flash284B-A13BQ4_K_MM3 Ultra256 GBMLXestimated32—2026-08
Inkling-Small276B-A12BQ4_K_MM3 Ultra512 GBMLXestimated35—2026-08
Qwen3-235B-A22B235B-A22BQ4_K_MM3 Ultra256 GBMLXestimated20—2026-08
Hunyuan Hy3295B-A21BQ4_K_MM3 Ultra256 GBLM Studioestimated21—2026-08
Qwen3-Coder-Next80B-A3BQ4_K_MM5 Max64 GBMLXestimated84—2026-08
Qwen3-Coder-Next80B-A3BQ4_K_MM4 Max64 GBMLXestimated77—2026-08
Qwen3.8-27B27.8BQ4_K_MM624 GBMLXestimated8—2026-08
Qwen 3.6-27B27BQ4_K_MM624 GBMLXestimated9—2026-08
KAT-Coder-V2.535BQ4_K_MM632 GBMLXestimated56—2026-08
DeepSeek R1 8B8BQ4_K_MM616 GBMLXestimated28—2026-08
Qwen 3.5 9B9BQ4_K_MM616 GBMLXestimated25—2026-08
Phi-4 14B14BQ4_K_MM616 GBMLXestimated16—2026-08
Gemma 4 E4B4BQ4_K_MM616 GBMLXestimated54—2026-08
LFM2.5-2.6B2.6BQ4_K_MM616 GBMLXestimated79—2026-08
DeepSeek V4 Flash284B-A13BQ4_K_MM5 Ultra256 GBMLXestimated45—2026-08
Inkling-Small276BQ4_K_MM5 Ultra256 GBMLXestimated49—2026-08
Qwen3-235B-A22B235BQ4_K_MM5 Ultra256 GBMLXestimated29—2026-08
Mistral Small 4119BQ4_K_MM5 Ultra96 GBMLXestimated79—2026-08
GPT-oss 120B117B-A5.1BMXFP4M5 Ultra96 GBMLXestimated97—2026-08
Llama 3.3 70B70BQ4_K_MM5 Ultra96 GBMLXestimated23—2026-08
DeepSeek R1 70B70BQ4_K_MM5 Ultra96 GBMLXestimated23—2026-08
Qwen3-Coder-Next80BQ4_K_MM5 Ultra96 GBMLXestimated130—2026-08
GLM-4.5-Air106BQ4_K_MM5 Ultra96 GBMLXestimated49—2026-08
Laguna S 2.1118BQ4_K_MM5 Ultra96 GBMLXestimated68—2026-08
Llama 3.1 405B405BQ4_K_MM5 Ultra512 GBMLXestimated4—2026-08
Qwen 2.5 72B72BQ4_K_MM5 Ultra96 GBMLXestimated23—2026-08
GLM-5.3-Flash320B-A18BIQ3_XXSM5 Max128 GBMLXestimated27—2026-09
GLM-5.3-Flash320B-A18BQ4_K_MM3 Ultra256 GBMLXestimated24—2026-09
GLM-5.3-Flash320B-A18BQ4_K_MM5 Ultra256 GBMLXestimated35—2026-09
Qwen3.8-Flash-Next125B-A6BQ4_K_MM5 Max128 GBMLXestimated49—2026-09
Qwen3.8-Flash-Next125B-A6BQ4_K_MM4 Max128 GBMLXestimated44—2026-09
Qwen3.8-Flash-Next125B-A6BQ4_K_MM3 Ultra256 GBMLXestimated62—2026-09
Qwen3.8-Flash-Next125B-A6BQ4_K_MM5 Ultra256 GBMLXestimated84—2026-09
DeepSeek V4 Flash Vision-Exp305B-A13BQ2_KM5 Max128 GBllama.cppestimated39—2026-09
DeepSeek V4 Flash Vision-Exp305B-A13BQ4_K_MM3 Ultra256 GBllama.cppestimated32—2026-09
DeepSeek V4 Flash Vision-Exp305B-A13BQ4_K_MM5 Ultra256 GBllama.cppestimated46—2026-09
Qwen3.8-27B27.8BQ4_K_MM5 Max128 GBMLXestimated29—2026-09
Qwen3.8-27B27.8BQ4_K_MM4 Max64 GBMLXestimated26—2026-09
Qwen3.8-27B27.8BQ4_K_MM3 Max64 GBMLXestimated19—2026-09
Qwen3.8-27B27.8BQ4_K_MM4 Pro48 GBOllamaestimated13—2026-09
GLM-5.3753B-A40BQ2_KM5 Ultra256 GBMLXestimated27—2026-09
GLM-5.3753B-A40BQ2_KM3 Ultra256 GBMLXestimated19—2026-09
GLM-5.3753B-A40BQ4_K_MM5 Ultra512 GBMLXestimated17—2026-09
GLM-5.3753B-A40BQ4_K_MM3 Ultra512 GBMLXestimated11—2026-09
Qwen3.8-27B27.8BQ4_K_MM5 Ultra256 GBLM Studiosourced550.5s2026-09
Qwen3.8-27B27.8BQ4_K_MM3 Ultra96 GBLM Studiosourced40—2026-09
Qwen 3.5 122B-A10B122B-A10BunstatedM5 Ultra256 GBLM Studiosourced80—2026-09
Qwen 3.5 122B-A10B122B-A10BunstatedM3 Ultra96 GBLM Studiosourced60—2026-09
Qwen 3.5 122B-A10B122B-A10BQ4_K_MM5 Max128 GBMLXestimated47—2026-09
Qwen 3.5 122B-A10B122B-A10BQ4_K_MM4 Max128 GBMLXestimated42—2026-09
Qwen 3.5 122B-A10B122B-A10BQ4_K_MM3 Max128 GBMLXestimated32—2026-09
MiMo-V2.6-Flash309B-A15BMXFP4M3 Ultra256 GBMLXestimated30—2026-09
MiMo-V2.6-Flash309B-A15BMXFP4M5 Ultra256 GBMLXestimated42—2026-09
Qwen 3.6-35B-A3B35B-A3B4-bitM5 Max128 GBMLXcommunity126—2026-08
Qwen 3.6-35B-A3B35B-A3B4-bitM632 GBMLXcommunity64—2026-10
Nemotron 3.5 Lightning30B-A3BQ4_K_MM632 GBllama.cppcommunity45—2026-10
Gemma 4 26B-A4B26B-A4BQATM632 GBLM Studiocommunity48—2026-10
Gemma 4 31B31BQATM632 GBLM Studiocommunity8—2026-10
Muse Glimmer 30B30BunstatedM632 GBLM Studiocommunity9—2026-10
GPT-oss 20B21B-A3.6BunstatedM632 GBLM Studiocommunity45—2026-10
Gemma 4 12B12BQATM632 GBLM Studiocommunity19—2026-10
Qwen 3 30B-A3B30B-A3B4-bitM4 Max128 GBMLXsourced110—2026-01
Qwen 3 30B-A3B30B-A3B4-bitM4 Max128 GBllama.cppsourced90—2026-01
GPT-oss 20B21B-A3.6BMXFP4M5 Max64 GBMLXestimated103—2026-10
GPT-oss 20B21B-A3.6BMXFP4M4 Max64 GBMLXestimated97—2026-10
GPT-oss 20B21B-A3.6BMXFP4M4 Pro24 GBMLXestimated64—2026-10
GPT-oss 20B21B-A3.6BMXFP4M416 GBMLXestimated34—2026-10
Gemma 4 12B12BQ4_K_MM5 Max64 GBMLXestimated64—2026-10
Gemma 4 12B12BQ4_K_MM4 Max64 GBMLXestimated57—2026-10
Gemma 4 12B12BQ4_K_MM4 Pro24 GBMLXestimated30—2026-10
Gemma 4 12B12BQ4_K_MM416 GBMLXestimated13—2026-10
MiniMax M2.5230B-A10BQ4_K_MM5 Ultra256 GBMLXestimated57—2026-10
MiniMax M2.5230B-A10BQ4_K_MM3 Ultra256 GBMLXestimated41—2026-10

Methodology

According to the LLMCheck index, all benchmarks measure tokens per second (tok/s) during the generation phase, excluding prompt processing time. This reflects the sustained output speed you experience when the model is actively generating text.

Time to first token (TTFT) is measured separately in seconds — the delay between submitting your prompt and receiving the first output token. TTFT depends on prompt length, model size, and available memory bandwidth.

Unless noted otherwise, figures assume Q4_K_M quantization (4-bit with k-quant medium), the most popular level for balancing quality and speed. The reference protocol for a measured row is a standardized 256-token prompt generating 512 tokens at default context settings, averaged over 3 runs on an otherwise idle machine; that is what vendors and contributors are asked to report. Estimated rows are not measured at all — they are computed from the model's memory footprint and the chip's bandwidth. Check the mark on each speed cell.

LLMCheck benchmarks are sourced from community submissions and verified against known baselines. Chip names refer to the full SoC variant (e.g., "M4 Pro" means the M4 Pro chip specifically, not the base M4). RAM indicates the total unified memory of the test system.

Next step
You have the tok/s figure. The two questions it usually raises next:
Already own the Mac — see every model that fits your exact chip and RAM, ranked on these same figures →
Still choosing one — the Mac Advisor turns a budget into a specific config →
Comparing the machines themselves? The Mac hardware guide sets out memory bandwidth, RAM tiers and prices side by side.
On an Intel Mac? These figures are Apple Silicon only — what an Intel Mac can realistically run, and how fast.

Frequently Asked Questions

How are these benchmarks measured?

Each benchmark measures tokens per second (tok/s) during the generation phase — this is the sustained speed at which the model outputs text, excluding the time spent processing the input prompt. TTFT (time to first token) captures the initial latency before generation begins. The reference protocol for a measured row is a standardized 256-token input prompt, 512 output tokens, Q4_K_M quantization and default context settings, averaged over 3 consecutive runs on an otherwise idle machine — what vendors and contributors are asked to report. Estimated rows are computed from memory footprint and bandwidth instead; the mark on each speed cell says which you are looking at.

Why does tok/s vary between Ollama, LM Studio, and MLX?

Each engine uses a different inference backend with distinct optimizations. MLX is Apple's native framework, purpose-built for Metal GPU acceleration on Apple Silicon — it often delivers the fastest results, especially for smaller models. Ollama uses llama.cpp with Metal support and provides reliable, consistent performance. LM Studio also wraps llama.cpp but adds a GUI layer that can introduce minor overhead. The performance gap between engines is typically 5-15% for the same model and hardware configuration.

Which Apple Silicon chip is best for local AI?

It depends on your target model size. For small models (3-9B), even an M1 with 16 GB delivers usable speeds (40-80 tok/s). For mid-size models (14-35B), the M4 Pro with 24 GB is the sweet spot — enough RAM for 14B models at 35-55 tok/s. For large models (70B+), the M5 Max with 128 GB is ideal, offering 614 GB/s memory bandwidth. The Mac Studio M5 Ultra (1.2 TB/s) and M3 Ultra (819 GB/s), with up to 512 GB, handle the biggest models but are overkill for anything under 70B.

Can I submit my own benchmarks?

Yes, we welcome community submissions. Run your benchmark using Ollama, LM Studio, or MLX with standard settings (Q4_K_M quantization, default context). Record your chip model, total RAM, engine version, and both tok/s and TTFT values. Submit via our GitHub repository or by email. We verify all submissions against known performance baselines before adding them to the database.

What is the fastest local LLM on Apple Silicon?

According to the LLMCheck index as of August 2026, Maple Preview 20B-A1B is the fastest entry at 281 tokens per second on an M5 Pro (vendor-reported), with LFM2.5-2.6B at 220 tok/s on M5 Max (vendor-reported) and Gemma 4 E2B the fastest pure-estimate entry at ~261 tok/s. Among larger models, Qwen 3.6-27B (77.2% SWE-bench Verified) generates ~30 tok/s estimated on an M5 Max, and DeepSeek V4 Flash (284B-A13B MoE) achieves ~39 tok/s community-reported on a 128 GB M5 Max at 2-bit.

Why is memory bandwidth important for running AI on Mac?

Memory bandwidth determines how fast your Mac can feed model weights to the GPU during inference. The LLMCheck index shows a near-linear relationship: the M5 Max (614 GB/s bandwidth) generates tokens roughly 3x faster than an M2 Pro (200 GB/s). This is why Unified Memory architecture gives Apple Silicon an advantage — there's no CPU-to-GPU transfer bottleneck.

How does LLMCheck calculate its composite score?

The LLMCheck Score is a 0–100 composite metric: 50 points for model capability (sourced from Arena AI ELO, MMLU, and coding benchmarks), 25 points for Mac-specific speed (tok/s on M5 Max), 15 points for accessibility (minimum RAM), and 10 points for license openness. Full formula and per-model sources at /methodology.html.

Open data

Every measurement on this page in CSV and JSON, free under CC BY 4.0 — including the provenance field, so you can filter to sourced rows only.

Download the dataset →