Changelog
LLMCheck is updated monthly. Here's the running log of catalog additions, new benchmark data, and site changes — so "updated monthly" is something you can verify, not just a claim.
5 October 2026 — MoE refit, the M6 measured, and corrections to the hardware data
The first Mac mini M6 measurements, and a refit. A community benchmark of a 32 GB Mac mini M6 (raw data published) and an M5 Max run of Qwen 3.6-35B-A3B showed small mixture-of-experts models decoding two to four times faster than the index estimated; dense models matched within a token per second. Small MoEs (40B total or less) now use the two-term model with a fitted ~5 ms per-token overhead (six runs, three chips, mean error ~12%); large MoEs keep the conservative 0.42 factor; a model’s own measurement anchors its other chips. KAT-Coder-V2.5 (82) and Qwen 3.6-35B-A3B (81) now lead the leaderboard. Estimated rows no longer carry a time-to-first-token, which the estimator never modelled. Details on /methodology and in the October edition.
Catalog 91 models, 244 speed figures (11 published, 13 individual runs, 220 estimated; data version 2026-10-05). Added, verified on their Hugging Face cards: GPT-oss 20B (OpenAI, Apache 2.0, 21B-A3.6B) and Gemma 4 12B (Google, Apache 2.0). GPT-oss 120B had been costed as a dense model; it is a 117B MoE with 5.1B active, and its figures are corrected. MiniMax M2.5 gains estimates on the Ultras it actually fits.
Corrections. (1) Apple’s M5 bandwidths: the MacBook Pro configurations carried the M5 Pro at 273 GB/s and the M5 Max at 600 (Apple: 307 and 614), and those values fed the estimator. The efficiency constant was refit (0.801 → 0.783), so M5 Max estimates are unchanged, M5 Pro estimates rose up to 12% and other chips moved about 2%; binned configurations (36 GB M5 Max at 460 GB/s, 16 GB M6 at 153) now use their own figures. Sixteen pages quoting the old figures, and several quoting the base M3 at 200 GB/s (Apple: 100), are corrected with dated notes. (2) There is no M4 Ultra. Apple never made one, but the index listed a “Mac Studio M4 Ultra” with 18 estimated rows, two Best-by-Mac pages, Advisor configurations and Amazon links. All removed; its URLs redirect to the M5 Ultra, and about twenty articles that cited it carry corrections. (3) The Mac Advisor scaled every model as if the index’s reference speeds were M3 Max figures, overstating M5 Max and Ultra configurations by 1.5–2×; it now reads the same per-Mac figures as the Best-by-Mac pages, with provenance. (4) From 23 September the leaderboard ranked MiMo-V2.6-Flash #1 on an M5 Ultra estimate shown in its M5 Max column; models too big for any M5 Max are now ranked on the M5 Ultra and tagged. (5) Seven articles attributed figures to tests LLMCheck never ran, and an M5 comparison presented an impossible table as its own Ollama measurements; LLMCheck runs no tests of its own, and the wording and tables are corrected. (6) /hardware, the homepage picks, and two M5 guides now read their speeds from the index at build time instead of hand-copying them.
Guards added to the build: every chip must be one Apple made; M5 bandwidths must match Apple’s; no page may sell or describe an M4 Ultra; no estimated row may carry a TTFT; no MoE estimate may beat its active-weight bandwidth ceiling; estimates must rise with bandwidth; the compare tool must match the catalog; the Associates tag may appear only in the redirect Worker.
23 September 2026 — the first measurements, four models, an estimator fix
The Macs shipped and the first reviews are in. BGR measured Qwen3.8-27B at ~55 tok/s on an M5 Ultra and ~40 on an M3 Ultra (estimates: 57 and 39) and Qwen 3.5 122B-A10B at ~80 and ~60. Those four figures enter the dataset as published measurements; the M5 Ultra estimate for Qwen3.8-27B is replaced. MacStories’ figures used multi-token prediction on long prompts and stay out of the plain-decode dataset. Write-up: how the estimates held up.
Catalog 89 models, 258 speed figures (9 published, 5 individual runs, 244 estimated; data version 2026-09-23). Added, each verified on its Hugging Face card: Qwen 3.5 122B-A10B (Apache 2.0, February 2026, missed until now), MiMo-V2.6-Flash (Xiaomi, 309B-A15B, MIT, omnimodal; 192 GB Mac Studio), MiMo-V2.6-Pro (1.02T-A42B, MIT, server-class) and DeepSeek V4.1 Flash (552B encoder-decoder, MIT — no Mac runtime yet). Details. Muse Spark: still no weights on Meta’s Hugging Face organization.
Integrity fixes. Provenance marks now mean one thing everywhere (filled square = published figure, vendor or press; hollow circle = one person’s run). Best-by-Mac pages mark every speed — estimates read from the dataset had rendered as bare numbers — and fill a chip without its own row from the model’s nearest row, fixing cases like an M2 Ultra outrunning an M5 Ultra. Nine rows labelled Q8_0 carried Q4 speeds above their own bandwidth ceiling; the estimator now costs each row at its labelled quantization and the leaderboard reference stays at 4-bit. Present-tense “we benchmark” wording removed from 44 generated pages and a handful of posts; the March M5 Max guide’s results table, presented as our own tests, now carries an editor’s note and the index’s estimates. /software adds oMLX.
7 September 2026 — September edition, four new models, MoE speeds recalibrated
Catalog 85 models, 248 speed figures (data version 2026-09-07). Added, each verified on its Hugging Face card: GLM-5.3 (Zhipu, 753B-A40B, weights 25 Aug under the custom GLM-5.3 License — MIT-like below a US$10B MaaS gate; #8), GLM-5.3-Flash (320B-A18B, MIT, the “Ox Alpha” model, 26 Aug; #10), Qwen3.8-Flash-Next (125B-A6B Qwen4 preview, Qwen Community 1.0 — coding-assistant and MaaS products need a separate license; its 4-bit MLX is 112 GB, a 128 GB-tier model; #13) and DeepSeek V4 Flash Vision-Exp (305B-A13B, MIT, experimental, weights 31 Aug; #3). Qwen3.8-27B is the new #1 (Score 71), as August’s edition predicted; its entry now carries the card’s SWE-Bench Pro 61.7 (vendor scaffold) and Terminal-Bench 2.1 73.0, plus the M5 Max row its description had quoted without one. Qwen3.8-Flash-Next is now the top pick on every 128 GB-and-up Best-by-Mac page.
MoE speed estimates recalibrated. Sparse models now use active-parameter bytes scaled by a factor of 0.42, fitted to the index’s one community-measured MoE row (DeepSeek V4 Flash, M5 Max, 39 tok/s); the existing MoE estimates already sat in the 0.24–0.53 band. One published figure changed: DeepSeek V4 Flash on M5 Ultra, ~80 → ~47 tok/s; the two 25 August Studio posts carry a dated note. Documented on /methodology. Three dense models that the compare tool mislabelled as MoE are corrected. September edition, plus model pages for GLM-5.3-Flash and Qwen3.8-Flash-Next.
Site structure, shipped 27 Aug – 7 Sep. /software now server-renders its 13 app cards (they were built client-side into empty containers, invisible to crawlers). /models/ consolidated: 92 pages render, the rest 301 or meta-refresh to the Best-by-Mac page for their chip. The glossary is 37 indexable term pages instead of one. Stale post-launch titles fixed sitewide; a nav overlap between 760 and 980 px fixed; /benchmarks no longer describes its estimated rows as measured runs. validate.py gained guards for each of these.
25 August 2026 — the M6 / M5 Ultra wave
Apple refreshed the desktops (newsroom, 25 Aug; ships 22 Sep): Mac mini with M6 (2nm, 153–170 GB/s, from $899) or M5 Pro (307 GB/s, from $1,699); Mac Studio with M5 Max (460–614 GB/s, from $2,499) or the quad-die M5 Ultra (1.2 TB/s, 96–512 GB, from $5,499). The index gained 12 Mac configs (44 total in the Advisor), 13 Best-by-Mac pages (44 total), 25 estimated benchmark rows for M6 and M5 Ultra (227 total, every one provenance-marked), and 25 new model×chip pages. Every new-machine speed is estimated from Apple’s published bandwidth via the two-term model — dense by formula, MoE scaled from each model’s existing reference row — and will be replaced by sourced figures as reviews land after 22 September. Launch coverage: the full analysis, the $899 M6 mini, the M5 Ultra.
15 August 2026 — speed model refitted, Qwen3.8-27B, click-tracked affiliate links
The speed model was refitted, and 134 estimates changed. An audit against our own published formula found that 63 estimated figures were physically impossible — they implied a model reading its weights faster than the memory bus can deliver them. Llama 3.1 8B was listed at 75 tok/s on an M4 whose bandwidth caps it at 26. The cause was a single efficiency multiplier, which cannot fit both a 30B model (genuinely bandwidth-bound) and a 3B one (dominated by fixed per-token cost). The estimator is now a two-term formula whose constants are solved from the two vendor-published Apple Silicon figures in the dataset, then validated against a third it had never seen: predicted 24.7 tok/s vs 24 measured, +3%. Full derivation on the methodology page. No measured figure was altered — sourced and community rows are never recomputed, and two that had been overwritten were restored. The benchmarks table also now carries a provenance mark on every speed cell; it previously shipped 202 numbers with none, which contradicted our own rule.
New entry. Qwen3.8-27B (Alibaba, Apache 2.0) went public on 14 August and enters the index at Score 71, narrowly ahead of Qwen 3.6-27B (70) — both are estimated at ~29–30 tok/s under the refitted model, so the capability score is what separates them. It is a 27.8B dense model with a native vision encoder and a 262K context window; at 4-bit it needs about 19 GB, so it fits a 24 GB Mac. It benchmarks above Qwen 3.6-27B on every published coding row. Nobody has published an Apple Silicon tok/s figure for it, so the speed cell is estimated from the bandwidth model (the 16 GB 4-bit build against 600 GB/s), never presented as a measurement. Catalog is now 81 verified models. GLM-5.3 was announced on 14 August but no weights exist; per the provenance rule it is not listed.
Affiliate links are now measurable. All 460 Amazon links moved from direct amazon.com hrefs to a first-party redirect at /go/az on our own Cloudflare Worker, which counts the click and then forwards. No third-party script, no cookie, no change to the privacy position. The Associates tag now lives in one place — the Worker — so it cannot be dropped by an edit, and the destination is rebuilt server-side from an allowlist, so the link cannot be pointed anywhere else. Every affiliate button now names Amazon in its own label.
Fewer dead ends. Benchmarks, Software, the leaderboard and both programmatic hubs previously had no route onward to the pages that answer “which Mac, then?” — each now ends with an explicit next step. robots.txt was also corrected: its per-crawler groups were Allow-only, which per RFC 9309 meant the disallow rules in the wildcard group were invisible to Googlebot, Bingbot and every AI crawler. Each group now carries them.
August 2026 — catalog audit & grand update
The correction. A full audit cross-checked every catalog entry against official sources (Hugging Face orgs, lab blogs, license files). Entries that could not be verified were removed: the “Qwen 4 / 4.1” family, “Llama 5” family, “Phi-5” family, “Gemma 4.5”, “DeepSeek R2/R3”, “Mistral Medium 4” & “Voyage” line, “Command R+ 2”, and “Grok 4 Open”. Misnamed entries were reconciled to their real counterparts (“GLM 5.2 Air” → GLM-4.5-Air; “Mistral Voyage 24B” → Mistral Small 3.2 24B; “DeepSeek R2” slot → DeepSeek V3.2-Speciale); licenses were corrected (Kimi K3 & K2.x are custom/Modified MIT, not plain MIT); affected pages carry correction notes and removed URLs redirect. Their benchmark rows were removed from the open dataset.
The update. Catalog now 80 verified models. Added the real July–August wave: Muse Glimmer 30B (Meta’s Apache 2.0 return), DeepSeek V4 Flash (runs on 128GB Macs), Inkling & Inkling-Small (Thinking Machines), Kimi K3 (weights, corrected license), Qwen3.8-2.4T-A95B, Qwen 3.6-27B (new verified Mac #1), Laguna S/XS 2.1, Nemotron 3.5 Lightning, KAT-Coder-V2.5, Ling-3.0 Flash, Hunyuan Hy3, Solar Open 2, Apertus 1.5, GLM-4.7-Flash, and the on-device wave (Bonsai 27B, Maple Preview, LFM2.5-2.6B, Nanbeige4.2-3B). Benchmarks: 202 rows (now labeled estimated / vendor-sourced / community per row) incl. new M1 Ultra, M2 Max and M3 Ultra coverage; harness annotations added to SWE-Bench Pro claims. Published the August State of Open-Source Local LLMs report and reissued July as a corrected edition.
July 2026
Expanded the leaderboard to 79 models, including several entries later removed by the August audit (see above). Launched Best LLM by Mac pages (ranked per chip + RAM). Published the July State of Open-Source Local LLMs report. Privacy hardening: removed all third-party analytics and fonts; added privacy & terms pages. Truth & provenance layer: every figure now labeled estimated/sourced/community with published estimation math; launched the Mac Advisor, Cite & Contribute pages and the open llmcheck-data repo.
June 2026
Published the June State of Open-Source Local LLMs report. Added new tok/s benchmark data across M4 and M5 chips and expanded the guides & troubleshooting hub.
May 2026
Published the May State of Open-Source Local LLMs report. Grand catalog refresh across model families and RAM tiers.
April 2026
Published the first monthly State of Open-Source Local LLMs report and expanded the Apple Silicon benchmark database.
See the current rankings
81 verified models ranked by speed, capability, RAM and license.
→ Open the Leaderboard