Where Llama 5 Actually Stands
As of August 2026, Meta has released no model called Llama 5 — not a 405B, not a 70B, not a preview. The current state of the Llama line is unchanged from spring: Llama 4 Maverick is the open flagship, and Meta's next-generation Llama — reported under the internal codename "Avocado" — is expected in 2027. Any listing, leaderboard entry, or article describing Llama 5 benchmark scores today is describing a model that has not shipped.
Meta's actual open-weights activity in 2026 went a different direction. Rather than extending the Llama series, Meta introduced a new brand: Muse. The closed, frontier-scale system is called Muse Spark; the open release distilled from it is Muse Glimmer 30B, which arrived in August 2026. It is Meta's first open-weights model published under Apache 2.0 — a genuine break from the Llama Community License that governed every previous Meta release.
Short version: skip the Llama 5 rumor cycle. The Meta model you can download, run, and build on today is Muse Glimmer 30B — and for Mac users it is arguably a better local fit than a hypothetical 405B ever would have been.
What Muse Glimmer 30B Is
Muse Glimmer 30B is a 30B-parameter dense multimodal agent model. Three design choices define it:
- Distilled from Muse Spark — the model is a distillation of Meta's closed frontier system, tuned specifically for agentic behavior: tool calls, multi-step plans, and computer-use-style tasks, rather than raw chat benchmarks.
- Natively multimodal — image input is supported in the base release and in the day-zero Ollama build, not bolted on via a separate vision adapter.
- Shipped with a drafter — the release bundles DFlash, a small speculative-decoding draft model that roughly doubles generation speed on Apple Silicon (details below).
Context window is 131K tokens, and at Q4 quantization the weights come to roughly 17–18 GB. That is the practical headline for Mac users: this is a Meta model aimed at the mid-range — a 24GB Mac loads it, a 32GB Mac runs it comfortably — not at 128GB workstations.
Benchmark Scorecard
Meta's vendor-reported numbers position Muse Glimmer as a strong generalist agent with exceptional math, rather than a coding specialist:
| Benchmark | Muse Glimmer 30B | What It Measures |
|---|---|---|
| AIME | 94.7 | Competition math |
| GPQA | 83.5 | Graduate-level science Q&A |
| SWE-Bench Pro | 51.2% | Hard real-world software engineering |
| Context window | 131K tokens | Long-document handling |
All scores above are vendor-reported. The 94.7 AIME and 83.5 GPQA are remarkable for a 30B dense model and reflect the Muse Spark distillation. The 51.2% SWE-Bench Pro is respectable — SWE-Bench Pro is a substantially harder harness than SWE-bench Verified — but it is not class-leading: on agentic-desktop and terminal evaluations (OSWorld, Terminal-Bench), Muse Glimmer trails Qwen 3.6-27B, the current LLMCheck Mac #1. Treat Muse Glimmer as the multimodal generalist, not the top coder.
Mac Performance & the DFlash Drafter
Meta published Apple Silicon figures at launch — itself a sign of where this model is aimed. The interesting part is the second column: DFlash, the bundled draft model, uses speculative decoding to nearly double effective generation speed. The drafter proposes several tokens ahead; the main model verifies them in a single pass, cutting the number of full 30B forward passes per token.
| Mac | Base speed | With DFlash | Provenance |
|---|---|---|---|
| M5 Max | 26.6 tok/s | 50.2 tok/s | Vendor-reported |
| M4 Max | 23.7 tok/s | 37.8 tok/s | Vendor-reported |
| M4 Pro (24GB) | ~10 tok/s | — | Community |
According to the LLMCheck index, expect around 27 tok/s on an M5 Max at Q4_K_M via MLX — closely matching Meta's base figure. The DFlash numbers are vendor-reported and workload-dependent: speculative decoding gains are largest on predictable text (code, structured output) and smaller on high-entropy creative writing. Even the base speeds clear comfortable reading pace on Max-class chips, and the 24GB M4 Pro community figure of ~10 tok/s makes it usable, if not brisk, on the smallest supported configuration. Full per-chip details are on the Muse Glimmer 30B benchmark page.
Setup on a Mac
Ecosystem support was unusually complete at launch: Ollama 0.32.x shipped a day-zero MLX build including image input, and SGLang's new MLX backend also runs it. The drafter is bundled — no separate download.
At Q4_K_M plan for ~18 GB of weights plus context. On 24GB Macs, keep other memory-hungry apps closed; on 32GB and up there is comfortable headroom for long contexts and image input.
Muse Glimmer vs Qwen 3.6-27B
The obvious comparison is Alibaba's Qwen 3.6-27B — the same weight class, the same Apache 2.0 license, and the current holder of the LLMCheck Mac #1 spot (Score 72).
| Metric | Muse Glimmer 30B | Qwen 3.6-27B |
|---|---|---|
| Coding (SWE-bench) | 51.2% SWE-Bench Pro (vendor) | 77.2% SWE-bench Verified |
| Math (AIME, vendor) | 94.7 | — |
| Multimodal | Yes — native image input | No — text only |
| Speed, M5 Max | 26.6 tok/s (50.2 w/ DFlash, vendor) | ~40 tok/s (estimated) |
| Context | 131K | 262K |
| License | Apache 2.0 | Apache 2.0 |
| Min RAM (Q4) | ~18 GB | ~18 GB |
The split is clean. For pure coding and terminal-agent work, Qwen 3.6-27B remains the stronger pick — higher verified coding scores, double the context, faster baseline generation. Muse Glimmer wins wherever images enter the workflow — screenshot-driven agents, document understanding, UI automation — and on math-heavy tasks. According to the LLMCheck index, Qwen keeps the overall Mac #1; Muse Glimmer slots in as the multimodal-agent pick for 24–32GB Macs.
The strongest open models of August 2026 — GLM 5.2, Kimi K3, Qwen3.8's open weights — are server-class: Kimi K3 and Qwen3.8 fit no Mac at all, and GLM 5.2 needs a 256GB M3 Ultra even at 1-bit (community-reported ~22 tok/s). To run them at full precision, rent a datacenter GPU by the minute on Vast.ai — often 5–6× cheaper than AWS or GCP, with H100s and B200s available on demand.
Vast.ai referral link — we may earn a small commission at no extra cost to you. It never influences our reviews or rankings.
License: Meta Goes Apache 2.0
Muse Glimmer 30B ships under Apache 2.0 — and for Meta, that is news in itself. Every previous Meta open release carried a Llama Community License: the 700M monthly-active-user clause, an acceptable-use policy layered over the grant, and "Built with Llama" branding requirements. None of that applies here.
- No MAU threshold — commercial use at any scale, no separate license to request.
- No branding requirement — no mandated attribution in products built on the model.
- OSI-approved — Apache 2.0 is genuine open source, matching the licensing of Qwen 3.6 and most of the current Mac-runnable field.
Whether this signals a permanent licensing shift for the eventual next-generation Llama is unknown. For Muse Glimmer specifically, the license question that dominated every prior Meta model review simply goes away.
The Verdict
Muse Glimmer 30B is the most Mac-relevant model Meta has ever released. It is not the best open coder — Qwen 3.6-27B holds that ground — and it is not frontier-scale. What it is: a genuinely multimodal, agent-tuned 30B under a clean Apache 2.0 license, with vendor-published Apple Silicon numbers, a bundled drafter that roughly doubles throughput, and day-zero Ollama support. That combination did not exist in the open ecosystem before August 2026.
Skip Muse Glimmer if…
Your workload is pure text coding. Qwen 3.6-27B posts 77.2% SWE-bench Verified at ~30 tok/s estimated on an M5 Max from the same ~18 GB footprint, with double the context window. For terminal agents and code review, it remains the LLMCheck Mac #1.
Run Muse Glimmer 30B if…
You want a local agent that can see. Screenshot-driven automation, document and image understanding, math-heavy reasoning, and multimodal chat on a 24–32GB Mac — all under Apache 2.0, at a vendor-reported 50.2 tok/s on M5 Max with DFlash enabled.
And on the question that brought most readers here: there is no Llama 5 to wait for this year. If the 2027 release lands, we will index it when the weights are public and verifiable. Until then, Meta's open-model story runs through Muse — and Glimmer is a strong opening move.