The 30-Day Recap (TL;DR)
Every major verified open-weights event of July 2026, with a one-line takeaway. The month had no single crown change on the Mac leaderboard — Qwen3.6-27B holds #1 — but it had more verified frontier activity than any month this year.
- GLM 5.2 since mid-Jun — Zhipu's 753B MoE (MIT) carries into July as the open-frontier headline: 68.5% SWE-Bench Pro on Zhipu's own scaffold, ~62% on Scale's standardized SEAL harness — the open lead on both.
- Inkling mid-Jul — Thinking Machines Lab's debut: Apache 2.0, 975B-A41B multimodal MoE, 77.6 SWE-bench Verified, Artificial Analysis Intelligence Index 41.
- Inkling-Small mid-Jul — 276B-A12B, Apache 2.0, 80.2% SWE-bench Verified — the open-weight record; 2-bit MLX (88.4 GB) fits a 128 GB Mac.
- Kimi K3 mid–late Jul — Moonshot's 2.8T MoE, the largest open-weight model ever (~1.56 TB); custom revenue-gated license; #1 open on LiveBench Coding; not Mac-runnable.
- Laguna XS / S 2.1 early / late Jul — Poolside's coding line under OpenMDW-1.1; S 2.1 (118B-A8B) hits 70.2 on Terminal-Bench 2.1 and fits 96–128 GB Macs at INT4 (~59 GB) with official MLX builds.
- Solar Open 2 250B late Jul — Upstage's 250B-A15B with 1M context under an Apache-based Solar license; no Mac quants yet.
- Apertus 1.5 8B / 70B late Jul — ETH+EPFL's fully open release — weights and training data — with 262K context and multimodal input.
- Hunyuan Hy3 early–mid Jul — Tencent's 295B-A21B under Apache 2.0, with official 1-bit and 4-bit GGUFs — a first from a major lab.
- DeepSeek V4-Flash-0731 end of Jul — MIT, 284B-A13B, Artificial Analysis Index 50 (top open tier); 2-bit MLX (96.5 GB) runs on a 128 GB Mac at ~39 tok/s on an M5 Max (community).
- The on-device wave throughout Jul — Bonsai 27B (ternary, 8 GB Macs), Nanbeige4.2-3B (looped-transformer), and KAT-Coder-V2.5 (35B-A3B) put real coding scores in tiny footprints — all vendor-reported for now.
Headline story: the frontier stayed open — and got crowded. GLM 5.2 kept the open-weights lead on SWE-Bench Pro under both harnesses, a brand-new lab took the SWE-bench Verified open record in its first month, and by the last day of July a top-tier open model fit on a 128 GB Mac for the first time. The Mac-runnable crown itself didn't move: Qwen3.6-27B holds it.
GLM 5.2 — the open frontier, harness-checked
Zhipu AI's GLM 5.2 — a 753-billion-parameter mixture-of-experts model under MIT, released in mid-June 2026 — remained the open-frontier reference point all through July. Its headline number needs a footnote, and the footnote matters: the widely quoted 68.5% on SWE-Bench Pro comes from Zhipu's own agent scaffold. On Scale's standardized SEAL harness, which runs every model through the same agent loop, the best open result is roughly 62% — and GLM 5.2 leads open-weights models there too. Whether it is "ahead of GPT-5 and Claude" depends entirely on which harness you trust; on the standardized one, the closed frontier still holds an edge. According to the LLMCheck index, both numbers are real, and neither should be quoted without its harness.
GLM 5.2 at a glance
Coding, by harness
The Mac story is real, but extreme. Community quantizations now squeeze GLM 5.2 onto the very top of the Mac Studio line: a 1-bit GGUF runs at roughly 22 tok/s on a 256 GB M3 Ultra, and a 4-bit MLX build reaches about 15 tok/s on a 512 GB M3 Ultra (both community-reported). That makes GLM 5.2 the largest model anyone has run on a single Mac — and still a poor practical choice for almost everyone.
The practical GLM on a Mac is GLM-4.5-Air. The 106B-A12B MoE (MIT) runs on a 64 GB Mac at an estimated ~30 tok/s at Q4-class quantization — see the GLM-4.5-Air model page. The newer lightweight sibling, GLM-4.7-Flash (31B dense, MIT), targets the 24–32 GB tier. The full GLM 5.2 story, including the harness question, is in our dedicated GLM 5.2 deep dive.
# Practical GLM on a Mac: GLM-4.5-Air (106B-A12B, MIT, 64 GB tier)
ollama run glm-4.5-air
# Lighter, newer: GLM-4.7-Flash (31B dense, MIT, 24-32 GB tier)
ollama run glm-4.7-flash
# Full GLM 5.2 (753B): community 1-bit GGUF, 256 GB M3 Ultra minimumVerdict: GLM 5.2 is the open-frontier reference model of mid-2026 — on both the vendor scaffold and the standardized harness. Quote its scores with the harness attached, and run GLM-4.5-Air, not the flagship, unless you own a 256 GB Ultra.
Inkling & Inkling-Small — Thinking Machines' debut
The most significant new entrant of July was Thinking Machines Lab, which debuted in mid-July with two Apache 2.0 releases. Inkling is a 975B-A41B multimodal MoE that lands at 41 on the Artificial Analysis Intelligence Index and 77.6 on SWE-bench Verified — a credible frontier entry on day one. The bigger story for this site is its sibling. Inkling-Small (276B-A12B) scores 80.2% on SWE-bench Verified — the highest score ever posted by an open-weights model — plus 89.5 on GPQA, and it fits on a Mac.
Inkling
Inkling-Small
MLX-only, for now. Inkling-Small ships with official MLX quantizations — the 2-bit build is 88.4 GB and runs on a 128 GB Mac Studio — while llama.cpp support is still pending, which means the ~130 GB GGUF IQ4 build for 192 GB machines exists but has no mainstream engine to run it yet. No reliable tok/s figures have surfaced; the LLMCheck index lists Inkling-Small with throughput marked "data pending" rather than an estimate.
# Inkling-Small, 2-bit MLX (88.4 GB) - 128 GB Macs, MLX only for now
mlx_lm.generate --model mlx-community/Inkling-Small-276B-A12B-2bit \
--prompt "Fix the failing tests in this repo"
# llama.cpp / Ollama support: pending as of August 2026Verdict: a brand-new lab took the open-weight SWE-bench Verified record in its first month, under Apache 2.0, with a build that fits a 128 GB Mac. If throughput lands anywhere usable once the tooling matures, Inkling-Small becomes the high-RAM Mac model to beat.
DeepSeek V4-Flash — top-tier weights on a 128 GB Mac
DeepSeek closed the month with V4-Flash-0731, released at the very end of July under MIT: a 284B-A13B MoE that lands at 50 on the Artificial Analysis Intelligence Index — the top open tier — and self-reports 82.7 on Terminal-Bench 2.1. The line that matters for this site: the 2-bit MLX build is 96.5 GB, which means a genuinely top-tier open model now runs on a 128 GB Mac.
The community made it fast. The strongest early numbers come from antirez's ds4 Metal engine: roughly 39 tok/s on an M5 Max and 27 tok/s on an M3 Max, both community-reported. In early August, llama.cpp merged DSpark speculative decoding for the V4 family, which should lift big-model decode speeds further. V4-Flash is not in Ollama yet — MLX and llama.cpp builds are the current paths. Per-chip details are on the V4-Flash model page.
# DeepSeek V4-Flash-0731, 2-bit MLX (96.5 GB) - 128 GB Macs
mlx_lm.generate --model mlx-community/DeepSeek-V4-Flash-0731-2bit \
--prompt "Refactor this module and add tests"
# Fastest path today: antirez's ds4 Metal engine (community)
# llama.cpp: DSpark speculative decoding merged early August
# Ollama: not yet supportedVerdict: V4-Flash is the July release with the most practical Mac impact. If you have a 128 GB machine, an Artificial-Analysis-top-tier open model at ~39 tok/s (community, M5 Max) is a category that simply did not exist a month earlier.
The server tier — Kimi K3 & Solar Open 2
Two more big July releases matter even though neither runs on a Mac. Kimi K3 launched mid-July with weights following in late July: a 2.8-trillion-parameter MoE weighing roughly 1.56 TB on disk — the largest open-weight model ever published. It ranks #1 among open models on LiveBench Coding and #3 overall on the Artificial Analysis Intelligence Index. The license is new territory: the custom "Kimi K3 License" is MIT-like but adds a commercial gate for model-as-a-service providers above $20M revenue. Solar Open 2 250B (Upstage, late July) is a 250B-A15B with a 1M-token context window under an Apache-based Solar license — a serious enterprise long-context play, with no Mac quantizations available yet.
| Kimi K3 | Solar Open 2 250B | GLM 5.2 | |
|---|---|---|---|
| Total size | 2.8T MoE (~1.56 TB) | 250B-A15B | 753B MoE |
| License | Kimi K3 License (revenue-gated) | Solar (Apache-based) | MIT |
| Coding signal | #1 open, LiveBench Coding | — | ~62% SWE-Bench Pro (SEAL std.) |
| Context | — | 1M | — |
| Mac status | Not runnable | No quants yet | 1-bit GGUF, 256 GB M3 Ultra (community) |
| Best for | Long-horizon coding agents (hosted) | Long-context enterprise pipelines | Open-frontier generalist |
Practical note: "open weights" and "runnable" have fully decoupled. K3's weights are public and no consumer machine on earth can load them. If you need frontier capability locally, the answers in July are V4-Flash or Inkling-Small on a 128 GB Mac — not the server tier.
Mac-runnable mid tier & the on-device wave
Below the frontier, July filled in every tier of the Mac range — from 96 GB Studios down to 8 GB Airs.
Laguna S 2.1 & XS 2.1 — Poolside goes open
Poolside shipped Laguna XS 2.1 (33B-A3B) in early July and Laguna S 2.1 (118B-A8B) in late July, both under the OpenMDW-1.1 license. S 2.1 posts 70.2 on Terminal-Bench 2.1 in max-thinking mode and 59.4 on the public SWE-Bench Pro set. The Mac fit is deliberate: INT4 comes in around 59 GB, targeting 96–128 GB machines, with official MLX builds and Ollama 0.32.x support. For terminal-centric agent work on a high-RAM Mac, it is the most interesting coding-specialist release of the month.
Hunyuan Hy3 — official 1-bit GGUFs
Tencent released Hunyuan Hy3 (295B-A21B, Apache 2.0) in early July and followed mid-month with something rarer: official 1-bit and 4-bit GGUFs. The 4-bit build needs a 192 GB machine; the 1-bit build brings a 295B-class model into 64 GB territory. A first-party lab shipping its own extreme quantizations — rather than leaving them to the community — is a precedent worth noting.
Apertus 1.5 — fully open, training data included
The Swiss AI initiative (ETH Zürich + EPFL) released Apertus 1.5 in 8B and 70B sizes in late July — fully open in the strictest sense, with training data published alongside the weights, 262K context, and multimodal input. Neither size tops a capability chart, but for regulated environments and researchers who need end-to-end provenance, Apertus is currently the only game in town at this scale.
The on-device wave — Bonsai, Nanbeige, KAT-Coder
Three July releases point at the same trend: serious capability in tiny footprints. Bonsai 27B (Prism ML, Apache 2.0) is a native ternary derivative of Qwen3.6-27B: a 3.9–5.9 GB checkpoint that runs on 8 GB Macs with roughly 90% quality retention, vendor-reported. Nanbeige4.2-3B (Apache 2.0) uses a looped-transformer architecture to squeeze a vendor-reported 63.6 SWE-bench Verified out of 3B parameters. KAT-Coder-V2.5 (Kwaipilot, Apache 2.0) is a 35B-A3B coding specialist on the Qwen3.6 base, reporting 69.4 SWE-bench Verified. All three numbers await independent replication — the LLMCheck index carries them with vendor-reported flags — but the direction is unmistakable.
Ecosystem: the tooling grew up too
Ollama 0.32.x turned typing ollama into an interactive coding agent, and its MLX engine matured meaningfully — DFlash draft-model support, M5 Neural Accelerator kernels, and image input. LM Studio 0.4.20 shipped alongside the new Bionic local-agent app. llama.cpp added Kimi K3 architecture support and, in early August, the DSpark speculative-decoding merge. The gap between "weights exist" and "weights run well on a Mac" is closing faster than at any point since MLX launched.
The full landscape — Open-Source Top 10 (July 2026)
According to the LLMCheck index, here is the open-source top 10 as of July 2026 (corrected). Score is the LLMCheck composite (capability + speed + accessibility + license, max 100) — it deliberately rewards models you can actually run, which is why record-setting server-class models sit below Mac-runnable ones. Mac Tier is the minimum unified memory for a comfortable quantized run.
| Rank | Model | Lab | Active | License | Mac Tier | Score |
|---|---|---|---|---|---|---|
| 1 | Qwen3.6-27B | Alibaba | 27B dense | Apache 2.0 | 24 GB | 72 |
| 2 | Qwen3.6-35B-A3B | Alibaba | 3B | Apache 2.0 | 24 GB | 70 |
| 3 | Inkling-Small | Thinking Machines | 12B | Apache 2.0 | 128 GB | 68 |
| 4 | DeepSeek V4-Flash-0731 | DeepSeek | 13B | MIT | 128 GB | 67 |
| 5 | Bonsai 27B | Prism ML | 27B ternary | Apache 2.0 | 8 GB | 66 |
| 6 | GLM-4.5-Air | Zhipu | 12B | MIT | 64 GB | 65 |
| 7 | KAT-Coder-V2.5 | Kwaipilot | 3B | Apache 2.0 | 24 GB | 64 |
| 8 | GLM 5.2 | Zhipu | MoE | MIT | 256 GB | 63 |
| 9 | Laguna S 2.1 | Poolside | 8B | OpenMDW-1.1 | 96 GB | 61 |
| 10 | Kimi K3 | Moonshot | MoE | Kimi K3 License | Server | 57 |
Three observations. First, the accessibility weighting is doing its job: Inkling-Small holds the raw capability record but ranks below Qwen3.6-27B, which nearly matches it on coding while running on a machine one-fifth the price. Second, Apache 2.0 and MIT still dominate, but two custom licenses cracked the top 10 — OpenMDW-1.1 (Laguna) and the revenue-gated Kimi K3 License — the first real license diversification in a year. Third, Thinking Machines debuted at #3 in its first month, the strongest first-month entry in the index's history. Hunyuan Hy3 and Solar Open 2 sit just outside the ten.
6 things that changed in July 2026
1. Every benchmark score now needs its harness attached
The GLM 5.2 gap — 68.5% on the vendor scaffold vs ~62% on Scale's standardized SEAL harness — is the clearest demonstration yet that agentic-coding scores are not comparable across harnesses. The same model, the same benchmark, a six-point spread. The LLMCheck index now annotates the harness on every SWE-Bench Pro figure, and this report quotes both numbers everywhere GLM 5.2 appears.
2. A new lab entered at the frontier — and a fully-open one entered at the base
Thinking Machines' Inkling pair proved a new entrant can land in the open frontier on day one, under Apache 2.0, with an open-weight record to its name. At the other end of the transparency spectrum, Apertus 1.5 shipped weights and training data — the strictest definition of open at meaningful scale this year. The field is widening in both directions at once.
3. "Open weights" no longer implies "you can run it"
Kimi K3's 1.56 TB checkpoint is public and unrunnable outside a datacenter; Solar Open 2 has no Mac quants at all. Meanwhile licensing is diversifying: OpenMDW-1.1, the Apache-based Solar license, and the revenue-gated Kimi K3 License all appeared in one month. Reading the license — and checking for a quant that fits your machine — now matters as much as the benchmark table.
4. The 128 GB Mac became the frontier tier
Two independent releases converged on the same target: Inkling-Small's 2-bit MLX build (88.4 GB) and V4-Flash's 2-bit MLX build (96.5 GB) both fit a 128 GB Mac and nothing smaller. A year ago the 128 GB tier was for comfortable 70B-class inference; as of July it is where genuinely top-tier open models live. Both releases were also MLX-first — llama.cpp support trailed or is still pending.
5. The ternary / on-device wave began
Bonsai 27B compressing a Qwen3.6-27B derivative into under 6 GB, Nanbeige4.2-3B's looped-transformer, and KAT-Coder-V2.5's 3B-active coding specialization are three different bets on the same thesis: most of the capability, a fraction of the footprint. All three carry vendor-reported numbers awaiting replication, but if even one holds up, the 8–16 GB Mac tier gets interesting again.
6. The toolchain went agentic and MLX-first
Ollama 0.32.x turned its CLI into a coding agent and matured its MLX engine (DFlash drafters, M5 Neural Accelerator kernels, image input); LM Studio 0.4.20 arrived with the Bionic local-agent app; llama.cpp added Kimi K3 support and merged DSpark speculative decoding in early August. Runtime improvements are now landing within days of model drops, not months.
By Mac tier — what to run TODAY (July 2026)
Updated recommendations as of July 2026, corrected. Speed figures are labeled by provenance — estimated (LLMCheck model), vendor-reported, or community — per the methodology. For a ranked list per exact Mac config, see Best LLM by Mac.
| Mac RAM | Primary pick | Speed / footprint | Backup |
|---|---|---|---|
| 8 GB | Bonsai 27B (ternary) | 3.9–5.9 GB checkpoint (vendor-reported) | Nanbeige4.2-3B |
| 16 GB | Qwen3.6-27B (Q3) | ~13 GB; est. low-30s tok/s | Gemma 4 12B |
| 24 GB | Qwen3.6-27B (Q4) | ~40 tok/s est. (M5 Max) | Qwen3.6-35B-A3B / KAT-Coder-V2.5 |
| 32–48 GB | Qwen3.6-27B (Q6/Q8) | est. | Gemma 4 31B / GLM-4.7-Flash |
| 64 GB | GLM-4.5-Air | ~30 tok/s est. | Hunyuan Hy3 (official 1-bit GGUF) |
| 96 GB | Laguna S 2.1 (INT4, ~59 GB) | official MLX; no tok/s published | GLM-4.5-Air (Q8) |
| 128 GB | DeepSeek V4-Flash (2-bit MLX) | ~39 tok/s community (M5 Max) | Inkling-Small (2-bit MLX) |
| 192 GB | Hunyuan Hy3 (official 4-bit GGUF) | — | Inkling-Small IQ4 (once llama.cpp lands) |
| 256 GB+ | GLM 5.2 (1-bit GGUF) | ~22 tok/s community (M3 Ultra) | GLM 5.2 4-bit MLX (512 GB, ~15 tok/s community) |
The 24 GB answer is unchanged: Qwen3.6-27B. 77.2% SWE-bench Verified, Apache 2.0, roughly 17 GB at Q4, and an estimated ~40 tok/s on an M5 Max — details on the Qwen3.6-27B model page. For faster drafting at some capability cost, its MoE sibling Qwen3.6-35B-A3B fires only 3B parameters per token.
The 128 GB tier is the change of the month. V4-Flash at ~39 tok/s (community, M5 Max) and Inkling-Small's record-holding weights both landed in the same tier within two weeks of each other. If you own a 128 GB Mac, July was the month it became a frontier machine.
What's coming next month
Speculative section — what LLMCheck is watching for in August 2026, based on public roadmaps and lab statements.
- Qwen3.8 open weights. Alibaba's hosted Qwen3.8-Max is imminent, and the company has publicly promised open weights for the 3.8 line. The open question is the license: the 3.6 line's Apache 2.0 is the industry's cleanest, and any tightening would be the story. A Mac-class 3.8 (a 27B successor) would reset the 24–32 GB tier overnight.
- DeepSeek V4 Pro GA. The GA of the 1.6T-class V4 Pro is expected in August. The April preview weights are MIT; whether the GA build's weights are published is the thing to watch. Either way it is server-class — the Mac story stays with V4-Flash.
- MLX engine momentum. Ollama's MLX engine (DFlash, M5 Neural Accelerator kernels) and llama.cpp's DSpark merge point at real tok/s lifts for exactly the big-MoE models that now fit 128 GB Macs. Expect the community V4-Flash and Inkling-Small numbers to move.
- llama.cpp support for Inkling-Small. The pending port is the gate on the open-record model reaching GGUF users and the 192 GB tier. When it lands, expect a fresh round of community throughput data.
Watch list: the single most consequential possible August event is the promised Qwen3.8 open-weights drop. If a Qwen3.8-27B arrives under a permissive license and clears Qwen3.6-27B's 77.2% SWE-bench Verified, the Mac-runnable crown moves; if the license tightens, Qwen3.6-27B stays the safe default and the story becomes the end of Apache-era Qwen.