Where Llama 5 Actually Stands

As of August 2026, Meta has released no model called Llama 5 — not a 405B, not a 70B, not a preview. The current state of the Llama line is unchanged from spring: Llama 4 Maverick is the open flagship, and Meta's next-generation Llama — reported under the internal codename "Avocado" — is expected in 2027. Any listing, leaderboard entry, or article describing Llama 5 benchmark scores today is describing a model that has not shipped.

Meta's actual open-weights activity in 2026 went a different direction. Rather than extending the Llama series, Meta introduced a new brand: Muse. The closed, frontier-scale system is called Muse Spark; the open release distilled from it is Muse Glimmer 30B, which arrived in August 2026. It is Meta's first open-weights model published under Apache 2.0 — a genuine break from the Llama Community License that governed every previous Meta release.

Short version: skip the Llama 5 rumor cycle. The Meta model you can download, run, and build on today is Muse Glimmer 30B — and for Mac users it is arguably a better local fit than a hypothetical 405B ever would have been.

What Muse Glimmer 30B Is

Muse Glimmer 30B is a 30B-parameter dense multimodal agent model. Three design choices define it:

Context window is 131K tokens, and at Q4 quantization the weights come to roughly 17–18 GB. That is the practical headline for Mac users: this is a Meta model aimed at the mid-range — a 24GB Mac loads it, a 32GB Mac runs it comfortably — not at 128GB workstations.

Benchmark Scorecard

Meta's vendor-reported numbers position Muse Glimmer as a strong generalist agent with exceptional math, rather than a coding specialist:

Benchmark Muse Glimmer 30B What It Measures
AIME 94.7 Competition math
GPQA 83.5 Graduate-level science Q&A
SWE-Bench Pro 51.2% Hard real-world software engineering
Context window 131K tokens Long-document handling

All scores above are vendor-reported. The 94.7 AIME and 83.5 GPQA are remarkable for a 30B dense model and reflect the Muse Spark distillation. The 51.2% SWE-Bench Pro is respectable — SWE-Bench Pro is a substantially harder harness than SWE-bench Verified — but it is not class-leading: on agentic-desktop and terminal evaluations (OSWorld, Terminal-Bench), Muse Glimmer trails Qwen 3.6-27B, the current LLMCheck Mac #1. Treat Muse Glimmer as the multimodal generalist, not the top coder.

Mac Performance & the DFlash Drafter

Meta published Apple Silicon figures at launch — itself a sign of where this model is aimed. The interesting part is the second column: DFlash, the bundled draft model, uses speculative decoding to nearly double effective generation speed. The drafter proposes several tokens ahead; the main model verifies them in a single pass, cutting the number of full 30B forward passes per token.

Mac Base speed With DFlash Provenance
M5 Max 26.6 tok/s 50.2 tok/s Vendor-reported
M4 Max 23.7 tok/s 37.8 tok/s Vendor-reported
M4 Pro (24GB) ~10 tok/s — Community

According to the LLMCheck index, expect around 27 tok/s on an M5 Max at Q4_K_M via MLX — closely matching Meta's base figure. The DFlash numbers are vendor-reported and workload-dependent: speculative decoding gains are largest on predictable text (code, structured output) and smaller on high-entropy creative writing. Even the base speeds clear comfortable reading pace on Max-class chips, and the 24GB M4 Pro community figure of ~10 tok/s makes it usable, if not brisk, on the smallest supported configuration. Full per-chip details are on the Muse Glimmer 30B benchmark page.

Setup on a Mac

Ecosystem support was unusually complete at launch: Ollama 0.32.x shipped a day-zero MLX build including image input, and SGLang's new MLX backend also runs it. The drafter is bundled — no separate download.

# Ollama 0.32.x — day-zero build with image input ollama run muse-glimmer-30b # Or MLX directly pip install mlx-lm mlx_lm.generate --model mlx-community/muse-glimmer-30b-q4_k_m --prompt "Hello!"

At Q4_K_M plan for ~18 GB of weights plus context. On 24GB Macs, keep other memory-hungry apps closed; on 32GB and up there is comfortable headroom for long contexts and image input.

Muse Glimmer vs Qwen 3.6-27B

The obvious comparison is Alibaba's Qwen 3.6-27B — the same weight class, the same Apache 2.0 license, and the current holder of the LLMCheck Mac #1 spot (Score 72).

Metric Muse Glimmer 30B Qwen 3.6-27B
Coding (SWE-bench) 51.2% SWE-Bench Pro (vendor) 77.2% SWE-bench Verified
Math (AIME, vendor) 94.7 —
Multimodal Yes — native image input No — text only
Speed, M5 Max 26.6 tok/s (50.2 w/ DFlash, vendor) ~40 tok/s (estimated)
Context 131K 262K
License Apache 2.0 Apache 2.0
Min RAM (Q4) ~18 GB ~18 GB

The split is clean. For pure coding and terminal-agent work, Qwen 3.6-27B remains the stronger pick — higher verified coding scores, double the context, faster baseline generation. Muse Glimmer wins wherever images enter the workflow — screenshot-driven agents, document understanding, UI automation — and on math-heavy tasks. According to the LLMCheck index, Qwen keeps the overall Mac #1; Muse Glimmer slots in as the multimodal-agent pick for 24–32GB Macs.

⚡ Need frontier scale? Rent a GPU

The strongest open models of August 2026 — GLM 5.2, Kimi K3, Qwen3.8's open weights — are server-class: Kimi K3 and Qwen3.8 fit no Mac at all, and GLM 5.2 needs a 256GB M3 Ultra even at 1-bit (community-reported ~22 tok/s). To run them at full precision, rent a datacenter GPU by the minute on Vast.ai — often 5–6× cheaper than AWS or GCP, with H100s and B200s available on demand.

Vast.ai referral link — we may earn a small commission at no extra cost to you. It never influences our reviews or rankings.

License: Meta Goes Apache 2.0

Muse Glimmer 30B ships under Apache 2.0 — and for Meta, that is news in itself. Every previous Meta open release carried a Llama Community License: the 700M monthly-active-user clause, an acceptable-use policy layered over the grant, and "Built with Llama" branding requirements. None of that applies here.

Whether this signals a permanent licensing shift for the eventual next-generation Llama is unknown. For Muse Glimmer specifically, the license question that dominated every prior Meta model review simply goes away.

The Verdict

Muse Glimmer 30B is the most Mac-relevant model Meta has ever released. It is not the best open coder — Qwen 3.6-27B holds that ground — and it is not frontier-scale. What it is: a genuinely multimodal, agent-tuned 30B under a clean Apache 2.0 license, with vendor-published Apple Silicon numbers, a bundled drafter that roughly doubles throughput, and day-zero Ollama support. That combination did not exist in the open ecosystem before August 2026.

Skip Muse Glimmer if…

Your workload is pure text coding. Qwen 3.6-27B posts 77.2% SWE-bench Verified at ~30 tok/s estimated on an M5 Max from the same ~18 GB footprint, with double the context window. For terminal agents and code review, it remains the LLMCheck Mac #1.

Run Muse Glimmer 30B if…

You want a local agent that can see. Screenshot-driven automation, document and image understanding, math-heavy reasoning, and multimodal chat on a 24–32GB Mac — all under Apache 2.0, at a vendor-reported 50.2 tok/s on M5 Max with DFlash enabled.

And on the question that brought most readers here: there is no Llama 5 to wait for this year. If the 2027 release lands, we will index it when the weights are public and verifiable. Until then, Meta's open-model story runs through Muse — and Glimmer is a strong opening move.