LLMCheck Index Methodology

How the LLMCheck index ranks 79 local and frontier LLMs for Mac — a published estimation model, linked third-party benchmarks, and community submissions. Fully transparent, fully reproducible.

LLMCheck is an independent index of local-LLM performance on Apple Silicon — not a benchmark lab. Every figure is either a transparent estimate from the published model below, a sourced third-party benchmark (linked per row), or a community submission. Each data point is labeled with its provenance, and no figure is claimed as a first-party lab measurement.

Where Every Number Comes From

Every row in the leaderboard, the benchmarks table, and the open dataset carries one of three provenance labels. This mirrors the provenance_policy field published in benchmarks.json:

Estimated

Derived from the LLMCheck estimation model — memory-bandwidth scaling plus quantization arithmetic, documented in full below. Estimates are useful for planning ("will this model fit and feel usable on my Mac?") but are not measurements. Every estimated speed figure on the site links back to this page.

Sourced

A published figure from an accountable source, with the URL attached to the row — a vendor’s model card or announcement, an independent press review that documents its setup, Arena AI ELO ratings, MMLU/HumanEval/SWE-Bench results from official model cards and papers, or independent hardware test data. We link, we don't re-host or re-run.

Community

One person’s run on a real Mac — a blog post, forum thread or benchmark submission — with its URL as the source. Community numbers are sanity-checked against the estimation model and known baselines before inclusion. Own a Mac? Submit a benchmark — real runs beat estimates every time.

Why an index and not a lab? A single lab machine can only test one chip, one runtime, one macOS version. An index that publishes its estimation math, links its sources, and accepts community runs covers the entire Apple Silicon range — and you can audit every step. When we're estimating, we say so.

How We Estimate Performance

Local LLM performance on Apple Silicon is unusually predictable, because two hardware numbers dominate everything: unified memory capacity (can the model fit?) and memory bandwidth (how fast can weights stream through the GPU?). The LLMCheck estimation model is built on those two numbers.

Step 1 — Will it fit? (RAM model)

model_size_GB (Q4_K_M) ≈ params_B × 0.57
≈ 4.5 bits/param effective — Q4_K_M is nominally 4-bit but carries scale factors, higher-precision layers, and embedding tables
+ KV cache: grows with context length — ~1–4 GB typical at 4k–32k context
+ macOS overhead: ~2–3 GB for the OS and the inference runtime
Fit rule: total must stay within ~75% of unified memory — macOS lets the GPU address roughly three-quarters of RAM by default

That 75% rule is where the minimum-RAM guidance across the index comes from: a 16 GB Mac has a ~12 GB working budget, a 24 GB Mac ~18 GB, a 64 GB Mac ~48 GB, and so on.

Step 2 — How fast? (Speed model)

Token generation is memory-bandwidth-bound: for every token, the runtime must stream essentially all active model weights from RAM through the GPU. Compute is rarely the bottleneck on Apple Silicon. That gives a simple ceiling:

est. tok/s = 1 ÷ ( model_bytes_GB ÷ (bandwidth_GB/s × 0.783) + 0.00146 )
bandwidth_GB/s: Apple’s published memory bandwidth for the chip (e.g. M4 Pro 273 GB/s, M5 Pro 307 GB/s, M5 Max 614 GB/s); binned parts such as the 32-core-GPU M5 Max (460 GB/s) or the 16 GB Mac mini M6 (153 GB/s) use their own figure
model_bytes_GB: weight bytes read per token — total size for dense models, active-parameter bytes for MoE
0.783: runtime efficiency on the bandwidth term — inside the ~0.6–0.8 range typical of llama.cpp / MLX-class runtimes. Mixture-of-experts models use the same form with their own constants (below)
0.00146: fixed per-token cost in seconds (~1.46 ms) — kernel launches and sampling, which do not shrink as the model does
Cross-chip scaling: tok/s scales with the bandwidth ratio only while the bandwidth term dominates — below roughly 4 GB of weights the fixed cost takes over, and doubling bandwidth stops doubling speed
Where those two constants come from — revised 15 August and 5 October 2026

They are solved, not chosen. The two unknowns are fixed by the two vendor-published Apple Silicon figures in this dataset, both on M5 Max: LFM2.5-2.6B at 220 tok/s and Muse Glimmer 30B at 27 tok/s.

The result was then checked against a figure it had never seen — Muse Glimmer 30B on M4 Max, vendor-published at 24 tok/s. The formula predicts 24.1, an error of +0.5%. The single community figure on an M4 Pro comes out about 23% high, which is what a thermally-limited laptop under sustained load should do.

5 October 2026: the fit was first solved with the M5 Max at 600 GB/s, the figure the index carried for its MacBook Pro configurations. Apple’s figure is 614 GB/s (and 307 GB/s for the M5 Pro, which the index had at 273), so the efficiency is 0.783 rather than 0.801. Bandwidth × efficiency on the M5 Max is unchanged, so every M5 Max estimate is unchanged; other chips’ estimates moved about 2% lower, M5 Pro estimates up to 12% higher, and the held-out M4 Max check improved from +3% to +0.5%.

One efficiency multiplier cannot fit both ends of the range. At 30B, decode really is bandwidth-bound; at 2–3B the fixed per-token cost dominates, and a pure bandwidth model predicts speeds no small model actually reaches. Every estimated figure in the index was recomputed with the formula above on 15 August 2026. Before that, 63 of them implied a model reading its weights faster than the memory bus could deliver them — not achievable on any hardware. Measured figures, the filled squares and hollow circles, were not touched by this and never are.

Mixture-of-experts models — refit 5 October 2026

A mixture-of-experts (MoE) model stores all of its parameters but reads only the active ones for each token, so its bandwidth term uses active-parameter bytes: a 35B MoE with 3B active reads about 1.7 GB per token instead of 20 GB. It also pays more fixed cost per token than a dense model — routing, gathering experts, more kernel launches per layer. Until October the index scaled the dense formula by a flat 0.42, fitted to a single 2-bit DeepSeek V4 Flash run. Six fully specified plain-decode measurements of small MoEs then came in at 2–4× those estimates:

ModelChipQuant / runtimeMeasuredSource
Qwen 3 30B-A3BM4 Max4-bit, MLX109.7 tok/sarXiv 2601.19139 (sourced)
Qwen 3 30B-A3BM4 Max4-bit, llama.cpp89.9 tok/sarXiv 2601.19139 (sourced)
Qwen 3.6-35B-A3BM5 Max4-bit, MLX, MTP off125.8 tok/soMLX database (community)
Qwen 3.6-35B-A3BMac mini M64-bit, MLX63.8 tok/sHDZucht, raw data (community)
Nemotron 3.5 LightningMac mini M6Q4_K_M45.1 tok/sHDZucht (community)
Gemma 4 26B-A4BMac mini M6QAT48.2 tok/sHDZucht (community)
small MoE (≤ 40B total): est. tok/s = 1 ÷ ( active_GB ÷ (bandwidth × 0.783) + 0.00497 )
large MoE (> 40B total): est. tok/s = 0.42 ÷ ( active_GB ÷ (bandwidth × 0.783) + 0.00146 )
0.00497: per-token overhead for small MoEs (~5 ms) — the median solved from the six runs above; mean absolute error ~12%, the dense model’s own accuracy class
0.42: the large-MoE factor, fitted to the only fully specified large-MoE runs (DeepSeek V4 Flash at 2-bit). Treat large-MoE estimates as floors: BGR’s Qwen 3.5 122B-A10B runs landed at 0.58–0.61 of the unscaled formula, quantization unstated
Own measurements first: a model with its own measured rows has every other chip scaled from its nearest-bandwidth measured row, using the same class model
Not bandwidth-bound: ternary and sub-2-bit formats decode through unpacking kernels; the index quotes only their measured figures and scales those, labelled estimated

The refit moved the small MoE models up the leaderboard: on an M5 Max several now estimate above 100 tok/s, where speed points cap at 25. The index’s #1, KAT-Coder-V2.5, ranks on an estimate from this model; #2, Qwen 3.6-35B-A3B, on its community-measured 126 tok/s. The full 35B of weights still has to be resident — about 20 GB — which is why it is a 24 GB-Mac model despite generating like a 3B one.

Step 3 — A worked example

▶ Llama 3.3 70B (dense) on an M5 Max, Q4_K_M

  1. Weights: 70 × 0.57 ≈ 40 GB on disk and in RAM.
  2. Total RAM needed: 40 + ~2 (KV cache) + ~3 (macOS) ≈ 45 GB.
  3. Fit check: a 64 GB Mac budgets 64 × 0.75 = 48 GB — fits. A 48 GB Mac budgets 36 GB — the 40 GB of weights alone don't fit. So a 64 GB Mac is the practical floor; the catalog records the model’s own need, about 44 GB.
  4. Speed ceiling: 614 GB/s ÷ 40 GB ≈ 15.4 tok/s theoretical maximum.
  5. Apply the formula: 1 ÷ (40 ÷ (614 × 0.783) + 0.00146) ≈ 12 tok/s — the estimated figure the index shows.

What the estimates can't capture

Honest limits of the model: thermals (a fanless MacBook Air throttles on long generations; a Mac Studio doesn't), context length (speed degrades as the KV cache grows — long chats get slower), quantization variants (Q5, Q8, and MLX 4-bit all shift size and speed), and runtime differences (Ollama, LM Studio, llama.cpp, and MLX can differ 10–20% on the same hardware). Real numbers will vary from the estimates — that's exactly why every estimated row is labeled, and why community submissions are invited to replace estimates with real runs.

Time to first token. The estimator models decode only. Since 5 October 2026 estimated rows carry no time-to-first-token; a TTFT appears only where a source published one. Prompt processing is where the M5 generation’s Neural Accelerators made their biggest gains, so a decode estimate says little about it.

The LLMCheck Score Formula

The estimation model above feeds one of four components in the composite 0–100 LLMCheck Score used to rank the catalog:

LLMCheck Score = Capability + Speed + Accessibility + License
Capability (0–50): normalized from Arena AI ELO + MMLU + coding benchmarks — sourced from published third-party evals
Speed (0–25): tok/s on the reference chip × 0.25, capped at 25 — measured where the index has a figure, otherwise estimated via the model above
Accessibility (0–15): RAM tier: ≤8 GB = 15, ≤16 GB = 12, ≤24 GB = 10, ≤32 GB = 9, ≤48 GB = 6, ≤64 GB = 5, above 64 GB = 3
License (0–10): MIT = 10, Modified MIT / GLM-5.3 = 9, Apache 2.0 / OpenMDW = 8, Gemma / Solar / LFM Open / NVIDIA Open = 6, Meta / Moonshot / Qwen Custom = 5, xAI Custom = 4, CC-BY-NC / N/A = 2

Dimension 1: Capability (50 points) — sourced

CAPABILITY50 / 100 pts

Raw model intelligence — reasoning, knowledge, coding, instruction following. Sourced from published third-party benchmark systems, linked per model.

The capability score (capScore) is the most heavily weighted dimension because the primary value of an LLM is output quality. capScore is sourced, not estimated: the index normalizes it from three public benchmark systems:

When official benchmarks are unavailable for a model (common for very new releases), the index extrapolates from related models in the same family and marks the score as estimated — same provenance rules as everything else. These estimates are replaced with sourced scores within 2–4 weeks of release.

Why 50 points? A very fast model that gives poor answers is less useful than a slower model with excellent reasoning. Capability is weighted highest because it determines whether the model actually solves your problem. Speed and RAM determine whether you can run it — but there's no point running a model that can't help you.

Capability Score Table (All 91 Models)

ModelParamscapScoreLicenseKey published evidenceSource
DeepSeek V4.1 Flash552B MoE50MITTerminal-Bench 2.1 90.6, DeepSWE v1.1 74.2, GPQA 90.9 (model card); no Mac runtime yetHF
MiMo-V2.6-Pro1.02T MoE50MITTerminal-Bench 2.1 89.9, DeepSWE v1.1 71.9, OSWorld-Verified 82.0 (model card)HF
MiMo-V2.6-Flash309B MoE49MITTerminal-Bench 2.1 87.6, DeepSWE v1.1 67.9 (model card)HF
Qwen 3.5 122B-A10B122B MoE38Apache 2.0SWE-bench Verified 72.0, GPQA 86.6, MMLU-Pro 86.7 (model card)HF
GLM-5.3753B MoE50GLM-5.3Terminal-Bench 2.1 88.2, DeepSWE v1.1 66.9 (model card); post-training upgrade of GLM-5.2HF
GLM-5.3-Flash320B MoE48MITTerminal-Bench 2.1 84.3 (model card); vendor claims above GLM-5.2 — no SWE-Bench Pro figure on the card as of 7 Sep 2026HF
Qwen3.8-Flash-Next125B MoE47Qwen CustomSWE-Bench Pro 62.5 (vendor scaffold), LiveCodeBench v6 91.9, GPQA 91.7 (model card)HF
DeepSeek V4 Flash Vision-Exp305B MoE47MITInherits V4 Flash text evidence; vision experimental, no independent benchmark verifiedHF
Qwen3.8-2.4T-A95B2.4T MoE50Qwen CustomSWE-Bench Pro 67.7, GPQA 92.6 (model card)HF
GLM 5.2753B MoE50MITSWE-Bench Pro 62.1% (SEAL standardized) / 68.5% vendor scaffoldHF
Kimi K32.8T MoE50Moonshot CustomAA Intelligence Index #3; Frontend Code Arena #1 (1,679)HF
DeepSeek V4 Pro1.6T MoE50MITSWE-bench Verified 80.6% (V4-Pro-Max, vendor)HF
Kimi K2.61T MoE48Modified MIT1T-A32B multimodal agent flagshipHF
GLM-5.1754B MoE48MITLed open SWE-Bench Pro before GLM 5.2HF
Inkling975B MoE47Apache 2.0AA Index 41; SWE-bench Verified 77.6HF
DeepSeek V4 Flash284B MoE47MITAA Index 50 (top open tier); Terminal-Bench 2.1 82.7 (vendor)HF
Kimi K2.51T MoE47Modified MIT—HF
Inkling-Small276B MoE46Apache 2.0SWE-bench Verified 80.2 (open record); GPQA 89.5HF
Qwen3-235B-A22B235B MoE46Apache 2.0—HF
DeepSeek V3.2-Speciale685B MoE45MITHigh-reasoning sibling of V3.2HF
Qwen 3.6-27B27B44Apache 2.0SWE-bench Verified 77.2HF
Laguna S 2.1118B MoE43OpenMDWTerminal-Bench 2.1 70.2 (max thinking); SWE-Pro public 59.4HF
Muse Glimmer 30B30B42Apache 2.0SWE-Bench Pro 51.2; AIME 94.7; GPQA 83.5 (vendor)HF
Solar Open 2 250B250B MoE42Solar LicenseMMLU-Pro 86.2; LiveCodeBench v6 92.4 (vendor)HF
DeepSeek V3.2685B MoE42MIT—HF
Hunyuan Hy3295B MoE41Apache 2.0AA Index 41HF
MiniMax M2.5230B MoE40Modified MITSWE-bench Verified 80.2 (launch coverage)HF
Gemma 4 31B31B40Apache 2.0—HF
KAT-Coder-V2.535B MoE39Apache 2.0SWE-bench Verified 69.4HF
Ling-3.0 Flash124B MoE38MITAA Index 38HF
Qwen 3.5 397B-A17B397B MoE38Apache 2.0—HF
Qwen 3.6-35B-A3B35B MoE38Apache 2.0SWE-bench 73.4HF
DeepSeek V3685B MoE37MIT—HF
DeepSeek R1671B MoE37MIT—HF
Nemotron 3.5 Lightning30B MoE36OpenMDWAA accuracy-vs-speed Pareto claim (vendor)HF
Mistral Large 3675B MoE36Apache 2.0—HF
Llama 4 Maverick400B MoE36Meta Custom—HF
GLM-4.7-Flash31B35MIT—HF
Llama 3.1 405B405B35Meta Custom—HF
Qwen3-Coder-Next80B MoE35Apache 2.0SWE-bench Verified 70.6HF
Gemma 4 26B-A4B26B MoE35Apache 2.0—HF
GLM-4.7355B34MIT—HF
Mistral Small 4119B MoE34Apache 2.0—HF
Laguna XS 2.133B MoE33OpenMDW—HF
Bonsai 27B27B33Apache 2.0~90% of Qwen3.6-27B FP16 quality (vendor)HF
GLM-4.5-Air106B MoE33MIT—HF
Step-3.5-Flash196B MoE33Apache 2.0—HF
MiMo-V2-Flash309B MoE32MIT—HF
GPT-oss 120B117B MoE32Apache 2.0—HF
Llama 4 Scout109B MoE30Meta Custom—HF
Nemotron-Cascade 230B MoE30NVIDIA OpenAIME / LiveCodeBench golds at 3B active (vendor)HF
Gemma 4 12B12B30Apache 2.0MMLU-Pro 77.2, GPQA Diamond 78.8, LiveCodeBench v6 72.0 (model card)HF
DeepSeek R1 70B70B30MIT—HF
Apertus 1.5 70B70B29Apache 2.0—HF
Hermes 4 70B70B28Meta Custom—HF
Mixtral 8x22B141B MoE28Apache 2.0—HF
Qwen 2.5 72B72B28Apache 2.0—HF
Llama 3.3 70B70B27Meta Custom—HF
Qwen 3.5 35B-A3B35B MoE27Apache 2.0—HF
QwQ 32B32B26Apache 2.0—HF
Mistral Small 3.2 24B24B25Apache 2.0—HF
Qwen 3 32B32B25Apache 2.0—HF
Devstral Small 24B24B24Apache 2.0—HF
DeepSeek R1 32B32B24MIT—HF
GPT-oss 20B21B MoE22Apache 2.0SWE-bench Verified 53.2, GPQA Diamond 58.6 at medium reasoning (model card)HF
Qwen 3 30B-A3B30B MoE22Apache 2.0—HF
Gemma 3 27B27B22Gemma—HF
Qwen 3.5 27B27B21Apache 2.0—HF
Maple Preview 20B-A1B20B MoE20MITIMO-level claims (vendor, unverified)HF
Nanbeige4.2-3B3B20Apache 2.0SWE-bench Verified 63.6 (vendor)HF
Qwen 3 14B14B20Apache 2.0—HF
Phi-4 14B14B19MIT—HF
Qwen 2.5 14B14B18Apache 2.0—HF
Ministral 3 14B14B18Apache 2.0—HF
Qwen 3.5 9B9B18Apache 2.0—HF
Gemma 3 12B12B17Gemma—HF
Apertus 1.5 8B8B16Apache 2.0—HF
Gemma 4 E4B4B16Apache 2.0—HF
DeepSeek R1 8B8B16MIT—HF
Qwen 3 8B8B15Apache 2.0—HF
LFM2.5-2.6B2.6B14LFM OpenSize-class instruction-following leader (vendor)HF
Ministral 8B8B14Apache 2.0—HF
Phi-4 Mini3.8B14MIT—HF
Gemma 4 E2B2B13Apache 2.0—HF
Mistral 7B7B13Apache 2.0—HF
Llama 3.1 8B8B12Meta Custom—HF
Qwen 3.5 4B4B12Apache 2.0—HF
Qwen 3 4B4B11Apache 2.0—HF
SmolLM3 3B3B10Apache 2.0—HF
Gemma 3 4B4B10Gemma—HF

Dimension 2: Speed on Apple Silicon (25 points) — estimated

SPEED25 / 100 pts

Tokens per second at 4-bit on the reference chip: the M5 Max — or, for models that need more than 128 GB, which no M5 Max can hold, the Mac Studio M5 Ultra, tagged on the leaderboard. A measured figure (sourced or community) is used where the index has one for that chip; otherwise the estimate from the model above. LLMCheck runs no lab.

Speed points are calculated as: est. tok/s × 0.25, capped at 25 points. Any model estimated at 100+ tok/s on the reference configuration receives full speed points. The estimates assume these reference conditions:

Models that fit no Mac (more than 512 GB) or have no Mac runtime yet receive 0 speed points. The full dataset — with a provenance label on every row — is available for download at /data/.

Dimension 3: Accessibility (15 points)

ACCESSIBILITY15 / 100 pts

How many Mac users can actually run this model? Lower RAM requirements = higher accessibility score.

Accessibility is a step function based on minimum RAM required at 4-bit quantization, derived from the RAM model above (weights + KV cache + macOS overhead, within the ~75% unified-memory budget). Since a large share of Mac users have 16 GB or less, accessibility is crucial for real-world impact:

Min RAM (4-bit)PointsExample ModelsEst. Mac Users
≤ 8 GB15Gemma 4 12B, Qwen 3.5 9B, Maple Preview 20B-A1B~100% of Apple Silicon
≤ 16 GB12GPT-oss 20B, Mistral Small 3.2 24B, Phi-4 14B~85%
≤ 24 GB10Qwen3.8-27B, Qwen 3.6-35B-A3B, KAT-Coder-V2.5~40%
≤ 32 GB9— (no current model)~30%
≤ 48 GB6Qwen3-Coder-Next, DeepSeek R1 70B, Apertus 1.5 70B~15%
≤ 64 GB5GLM-4.5-Air, GPT-oss 120B, Llama 4 Scout~10%
> 64 GB3DeepSeek V4 Flash, Qwen3.8-Flash-Next, GLM 5.2128 GB+ Macs, or server

Dimension 4: License Openness (10 points)

LICENSE10 / 100 pts

How freely can you use, modify, and distribute the model? More open = higher score.

LicensePointsCan Modify?Commercial Use?Models
MIT10YesUnrestrictedDeepSeek V4 Flash, GLM 5.2, Maple Preview
Modified MIT9YesYes (attribution clauses)Kimi K2.5/K2.6, MiniMax M2.5
Apache 2.08YesYes (with notice)Qwen 3.6, Muse Glimmer, Gemma 4, Inkling
OpenMDW-1.18YesYesNemotron 3.5 Lightning, Laguna S/XS 2.1
Gemma6YesYes (restrictions)Gemma 3 (old license)
Solar / LFM / NVIDIA Open6YesYes (attribution / terms)Solar Open 2, LFM2.5-2.6B, Nemotron-Cascade 2
Meta Custom5LimitedYes (<700M users)Llama 4, Llama 3.x, Hermes 4 70B
Moonshot / Qwen Custom5LimitedYes (revenue thresholds)Kimi K3, Qwen3.8-2.4T-A95B
CC-BY-NC2YesNon-commercial only(none currently listed)
Proprietary / N/A2NoAPI only(none currently listed)

Note: CC-BY-NC (non-commercial) scores 2 — usable for research but not commercial deployment. Meta / xAI community licenses score 5 / 4, reflecting modification rights with commercial caps.

Score Examples

Gemma 4 26B-A4B (Score: 67) = capScore 35 + speed min(25, est. 48 tok/s × 0.25 = 12) + accessibility 10 (24 GB) + license 8 (Apache 2.0) + rounding = 65–67. Top-ranked because it combines Arena AI #6 quality with fast MoE inference on a 24 GB Mac.

Qwen 3.5 9B (Score: 66) = capScore 18 + speed min(25, est. 100 tok/s × 0.25 = 25) + accessibility 15 (8 GB) + license 8 (Apache 2.0) = 66. Ranks high because maximum speed + accessibility points compensate for lower raw capability.

Kimi K2.5 (Score: 60) = capScore 50 (highest!) + speed 0 (server only) + accessibility 0 (>128 GB) + license 10 (MIT) = 60. Despite being the most capable model, it scores lower because no Mac user can run it locally.

Limitations & Known Issues

Corrections and real-world benchmark data are always welcome. If you have runs that differ from the index's estimates, submit them through the community benchmark process — verified community rows replace estimates.

Frequently Asked Questions

Does LLMCheck run its own benchmarks?

No. LLMCheck is an independent index, not a benchmark lab. Every figure on the site is one of three things: an estimate from the published LLMCheck estimation model (memory-bandwidth math and quantization arithmetic, fully documented on this page), a sourced number from a linked third-party benchmark, or a community-submitted run. Each row is labeled with its provenance, and no figure is claimed as a first-party lab measurement.

How does LLMCheck calculate its scores?

The LLMCheck Score is a 0–100 composite metric: Capability (50 pts) sourced from published third-party evaluations such as Arena AI ELO ratings and MMLU/coding benchmarks, Speed (25 pts) from estimated tokens/sec on the M5 Max reference configuration, Accessibility (15 pts) inversely proportional to minimum RAM, and License Openness (10 pts) where MIT scores 10 and restrictive licenses score lower. The formula is fully transparent and reproducible.

Where does LLMCheck get its capability scores?

Capability scores are derived from three public benchmark sources: Arena AI ELO ratings (human preference, weighted 40%), MMLU scores from official model cards (knowledge breadth, weighted 35%), and coding benchmarks like HumanEval and SWE-Bench (weighted 25%). All sources are linked per model. When official benchmarks are unavailable, the index extrapolates from related models in the same family and marks the score as 'estimated'.

How does LLMCheck estimate tokens per second on Apple Silicon?

Token generation on Apple Silicon is memory-bandwidth-bound, so estimated tok/s ≈ (memory bandwidth in GB/s ÷ model size in GB at Q4_K_M) × an efficiency factor of roughly 0.6–0.8 observed for llama.cpp/MLX-class runtimes. Mixture-of-Experts models use active-parameter bytes instead of total size, which is why they are much faster. The active-bytes figure is then scaled by a MoE factor of 0.42, fitted to the index’s one community-measured MoE row (DeepSeek V4 Flash, 13B active, M5 Max, 39 tok/s); routing, expert-loading and KV overheads mean a sparse model does not reach the bandwidth-bound ceiling its active size alone would suggest, and the existing MoE estimates all sit in the 0.24–0.53 band of the unscaled formula. Update, 23 September 2026: the first press measurements of a sparse model on the new Studios — Qwen 3.5 122B-A10B at ~80 tok/s on M5 Ultra and ~60 on M3 Ultra (BGR) — land near 0.6 of the unscaled formula. The index keeps 0.42 as a deliberately conservative factor rather than refit to a spread that wide, so MoE estimates may understate real speed; a model with its own measured rows has its other chips scaled from those measurements instead. On the Best-by-Mac pages, a chip with no row of its own is filled from the model’s nearest-bandwidth row using the same two-term model, and is always marked estimated. Scaling across chips follows the bandwidth ratio. The full model is published at llmcheck.net/methodology#estimation, and every estimated figure is labeled as such.

Why does LLMCheck weight capability at 50% of the total score?

Capability receives the highest weight because the primary value of an LLM is the quality of its outputs. A very fast model that gives poor answers is less useful than a slower model with excellent reasoning. However, speed (25%) and accessibility (15%) ensure that models which actually run well on consumer Macs score higher than server-only models with superior capability but no practical local use.

How often are LLMCheck scores updated?

Scores are updated within 48–72 hours of major model releases. The full leaderboard is refreshed monthly with the latest Arena AI ELO ratings and community benchmark submissions. Speed estimates are recomputed as new Apple Silicon hardware specifications become available. All updates are timestamped in the open dataset at llmcheck.net/data/.

Benchmark Sources

See the Full Leaderboard

91 verified models ranked by LLMCheck Score. Filter by your Mac's RAM, sort by speed or capability — every figure labeled with its provenance.

View Leaderboard →