What Happened
In July 2026 a data incident let several model names into the LLMCheck catalog that never existed as open releases — "Phi-5 Large 28B" among them. The name was plausible: Microsoft's phi family has iterated quickly, and a scaled-up synthetic-data 28B would be a natural next step. But plausible is not the same as real. Microsoft has announced no Phi-5 of any size, and no Phi-5 weights exist on Hugging Face, in Ollama's library, or anywhere else.
The catalog has since been reconciled against official vendor sources. Fabricated entries were removed from the leaderboard and benchmarks, and pages like this one carry corrections in place rather than quietly disappearing. Every speed figure on the site carries a provenance label — estimated, vendor-reported, or community — per the methodology.
Where Microsoft's Local Line Actually Stands
Microsoft's current local line is the Phi-4 family: Phi-4 14B and the smaller Phi-4 Mini. Phi-4 14B is the one that matters for this tier — a 14B dense model under the MIT license, built on the phi recipe of curated, largely synthetic "textbook-quality" training data.
It holds up well for its size: around 84.8% MMLU and unusually strong math for a 14B, which is the recipe's signature. At Q4 quantization the weights come to roughly 9 GB, so it fits a 16GB Mac with genuine headroom — still one of the better answers for that RAM tier. The trade-offs are the same ones the phi line has always had: it is text-only, its context window is a modest 16K tokens, and synthetic-data training leaves it thinner on long-tail world knowledge.
According to the LLMCheck index, Phi-4 14B remains a strong 16GB-tier pick — but at 24-32GB it is now outclassed by a wave of verified mid-2026 releases with bigger scores, bigger contexts, and in one case, vision.
What to Run Instead at 14-32GB
The tier the fictional "Phi-5 Large" supposedly owned is, in reality, the most contested segment in local AI right now. Three verified models lead it as of August 2026:
- Qwen3.6-27B — the LLMCheck Mac #1. A 27B dense model under Apache 2.0 that scores 77.2% on SWE-bench Verified — the strongest agentic-coding result of anything that fits a consumer Mac. It carries an LLMCheck Score of 72 and runs at roughly 40 tok/s (estimated) on an M5 Max at a ~16 GB Q4 footprint. If you want one model for a 24-32GB machine, this is it.
- Muse Glimmer 30B — the multimodal agent pick. Meta's August 2026 release under Apache 2.0, and its first open model since retiring the Llama open line. A 30B dense multimodal agent model distilled from the closed Muse Spark, it posts 51.2% SWE-Bench Pro, 94.7 AIME, and 83.5 GPQA. At Q4 it needs ~17-18 GB. Meta vendor-reports 26.6 tok/s on an M5 Max — 50.2 with the bundled DFlash drafter — and 23.7 (37.8 drafted) on an M4 Max, with day-zero Ollama MLX support including image input.
- Nemotron 3.5 Lightning 30B-A3B — the throughput play. NVIDIA's August 2026 release under OpenMDW-1.1 with fully open training data. It is a Mamba-hybrid MoE with ~3B active parameters, built for raw speed; the 4-bit MLX build is ~17.8 GB and ships as an official Ollama MLX build. No Apple Silicon tok/s figures have been published yet, so treat it as promising-but-unproven on Macs.
And if you are on 16GB, Phi-4 14B keeps its slot: nothing in the new wave undercuts its ~9 GB footprint while matching its reasoning.
Head-to-Head: The 24-32GB Tier
Here is how the field compares. Note the provenance on every speed figure — none of these numbers were produced by LLMCheck directly.
| Metric | Phi-4 14B | Qwen3.6-27B | Muse Glimmer 30B | Nemotron 3.5 Lightning |
|---|---|---|---|---|
| Architecture | 14B dense | 27B dense | 30B dense | 30B-A3B Mamba-hybrid MoE |
| License | MIT | Apache 2.0 | Apache 2.0 | OpenMDW-1.1 |
| Multimodal | No | No | Yes (image input) | No |
| Standout score | 84.8% MMLU | 77.2% SWE-bench Verified | 94.7 AIME / 83.5 GPQA | Not yet published |
| Q4 footprint | ~9 GB | ~16 GB | ~17-18 GB | ~17.8 GB (MLX 4-bit) |
| Fits | 16 GB Mac | 24-32 GB Mac | 24-32 GB Mac | 24-32 GB Mac |
| Speed (M5 Max) | ~75 tok/s (est.) | ~40 tok/s (est.) | 26.6 tok/s, 50.2 w/ DFlash (vendor) | Not yet published |
The shape of the choice: Qwen3.6-27B for coding and general work — its SWE-bench Verified lead is decisive and its estimated throughput is comfortable. Muse Glimmer 30B if you need vision or agentic workflows — its math and science scores are the best in the tier, and with the bundled DFlash speculative drafter its vendor-reported speed roughly doubles. Phi-4 14B if you have 16GB — half the footprint of everything else here. Nemotron 3.5 Lightning is the one to watch once community Apple Silicon numbers land; a 3B-active MoE should be very fast on paper.
One thing worth noting on benchmark hygiene: Muse Glimmer's coding score is on SWE-Bench Pro and Qwen's on SWE-bench Verified — different harnesses, not directly comparable. On the tasks where they meet (OSWorld, Terminal-Bench), Qwen3.6-27B stays ahead.
Setup Notes
All of the real models here are a short path from a fresh Mac. Phi-4 is one command in Ollama:
Qwen3.6-27B is available through Ollama and LM Studio in both GGUF and MLX builds. Muse Glimmer 30B and Nemotron 3.5 Lightning both shipped official Ollama MLX builds — Muse Glimmer's includes image input. Note that Ollama 0.32.x changed its default behavior: typing ollama with no arguments now opens an interactive coding agent, and the MLX engine picked up speculative decoding (DFlash) and M5 Neural Accelerator support — which is exactly what Muse Glimmer's doubled vendor-reported speeds rely on.
For editor integration, all of these serve an OpenAI-compatible endpoint at localhost:11434/v1, so Continue.dev and Cursor setups work unchanged — point them at the endpoint and set the model name to whichever you pulled.