What is Qwen 3.6, and which build should you run?
Qwen 3.6 is Alibaba's current open-weight generation for local use, released under the permissive Apache 2.0 license. Two builds matter on a Mac:
- Qwen3.6-27B — a dense 27B model and the quality pick. It scores 77.2% on SWE-bench Verified, and according to the LLMCheck index it is the current #1 Mac-runnable model with an LLMCheck Score of 72.[1]
- Qwen3.6-35B-A3B — a mixture-of-experts build with 35B total parameters but only ~3B active per token. It gives up some benchmark ground (73.4% SWE-bench) in exchange for materially faster generation and lighter compute per token.
| Spec | Qwen3.6-27B | Qwen3.6-35B-A3B |
|---|---|---|
| Architecture | Dense 27B | MoE, 35B total / ~3B active |
| License | Apache 2.0 | Apache 2.0 |
| SWE-bench | 77.2% (Verified) | 73.4% |
| RAM at Q4 | ~16 GB (18 GB+ Mac) | ~20 GB (24 GB+ Mac) |
| Speed on M5 Max (estimated) | ~40 tok/s | ~52 tok/s |
| LLMCheck Score | 72 — Mac #1 | See leaderboard |
The short version: run the 27B if you want the best answers your Mac can produce, and the 35B-A3B if you want the snappiest interactive feel on a 24 GB+ machine. Per-chip breakdowns are on the model pages for Qwen3.6-27B and Qwen3.6-35B-A3B.
Step 1: Check Your Memory & Pick Your Model
The single most important factor is unified memory. To check yours: click the Apple menu → About This Mac, and read the Memory line.
- 24 GB or more — both models are on the table. Pick based on the quality-vs-speed tradeoff above.
- 18 GB — run Qwen3.6-27B (~16 GB at Q4) with a modest context window; skip the 35B-A3B.
- 8–16 GB — neither fits comfortably at Q4. Run Bonsai 27B instead: Prism ML's Apache 2.0, natively ternary derivative of Qwen3.6-27B. Its checkpoint is 3.9–5.9 GB, it runs on 8 GB Macs, and it retains roughly 90% of the original's quality (vendor-reported). Details on the Bonsai 27B model page.
Shopping for a Mac to run Qwen 3.6? The sweet spot is an M4 Pro or M5-family chip with 24–32 GB of unified memory. Our Mac hardware buying hub breaks down which configuration gives the best tok/s per dollar for local LLMs — and the cheapest Mac that clears the 24 GB bar.
Step 2: Install Ollama
Ollama is the easiest way to run Qwen 3.6 on a Mac. Download it from ollama.com, open the .dmg, and drag Ollama into your Applications folder. Launch it once so it installs its command-line tool, then open Terminal and verify:
ollama --version
You should see ollama version 0.32.x or later — the 0.32 series matters because its MLX engine matured substantially on Apple Silicon, and typing plain ollama now opens an interactive coding agent. For the full walkthrough with screenshots and troubleshooting, follow our dedicated Install Ollama on Mac guide first, then come back here.
Step 3: Pull & Run Qwen 3.6
This is the part you came for. A single command downloads your chosen model and drops you into a chat:
# Quality pick — the LLMCheck Mac #1 (18 GB+ Macs)
ollama run qwen3.6:27b
# Speed pick — MoE, ~3B active (24 GB+ Macs)
ollama run qwen3.6:35b-a3b
The first run pulls the Q4 build (~16 GB for the 27B, ~20 GB for the 35B-A3B), so expect a wait on broadband. After that, the model is cached locally and launches in seconds. When you see the >>> prompt, you are talking to Qwen 3.6 entirely on your own machine — nothing leaves your Mac. To leave the chat, type /bye or press Ctrl+D. Want to download without starting a chat? Use ollama pull qwen3.6:27b.
According to the LLMCheck index, here is what to expect for generation speed (all figures estimated, Q4):
| Model | Mac | Speed (Q4, estimated) |
|---|---|---|
| Qwen3.6-35B-A3B | M5 Max | ~52 tok/s |
| Qwen3.6-27B | M5 Max | ~40 tok/s |
Anything above ~30 tok/s reads faster than most people, so both feel responsive on recent chips. Lower-end and older chips scale down from there — the Best by Mac pages list estimates for your exact configuration.
Step 4: Prefer an App? Run It in LM Studio
If you would rather skip the Terminal, LM Studio (free, 0.4.20 or later) runs the same models behind a polished graphical interface:
- Download LM Studio from lmstudio.ai and drag it into Applications.
- Open the Discover tab and search for
Qwen3.6. - Pick the MLX build of
Qwen3.6-27BorQwen3.6-35B-A3Bat a quantization that fits your RAM (LM Studio flags builds that will not fit). - Click Load, then chat in the Chat tab.
LM Studio also shows live tok/s and memory pressure, which makes it a nice way to sanity-check what your Mac can handle before committing to a workflow. Full setup details are in our LM Studio on Mac guide.
Step 5: Use It in Your Apps via the API
Ollama automatically serves an OpenAI-compatible API at http://localhost:11434, so any tool that speaks the OpenAI format can use Qwen 3.6 as a drop-in local backend. Here is a quick test with curl:
curl http://localhost:11434/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "qwen3.6:27b",
"messages": [{"role": "user", "content": "Write a haiku about Apple Silicon."}]
}'
The same endpoint works from Python with the official OpenAI SDK — just point base_url at localhost:
from openai import OpenAI
client = OpenAI(
base_url="http://localhost:11434/v1",
api_key="ollama" # required by the SDK, but ignored locally
)
resp = client.chat.completions.create(
model="qwen3.6:27b",
messages=[{"role": "user", "content": "Summarize the plot of Dune in 3 sentences."}]
)
print(resp.choices[0].message.content)
This makes Qwen 3.6 a free, private replacement for cloud APIs in scripts, agents, RAG pipelines, and editor extensions. (LM Studio exposes an equivalent local server on port 1234 if you went that route.)
Step 6: Tune It — Quantization, Context & MLX
Quantization (Q4 vs Q5/Q8)
The default tags are Q4 — the best balance of quality, speed, and memory for most Macs. If you have the headroom and want a small quality bump, pull a higher-precision build:
# Higher quality, more RAM and a bit slower
ollama pull qwen3.6:27b-q5_k_m
ollama pull qwen3.6:27b-q8_0
Q5 is a sensible step up on 32–48 GB Macs; Q8 is mainly for 64 GB+ machines where memory is not a constraint. Our quantization guide covers the tradeoffs in depth.
Context length
Longer contexts use more memory. To raise the window for a session, set num_ctx:
>>> /set parameter num_ctx 32768
On an 18–24 GB Mac, keep context modest (8K–16K) to avoid swapping. With 48 GB+ you can push the window much further for long-document work.
MLX for maximum speed
For the fastest possible inference on Apple Silicon, run Qwen 3.6 through MLX, Apple's native ML framework, directly — typically a step up from GGUF backends at the same quantization:
# one-time install
pip install mlx-lm
# generate from the command line
mlx_lm.generate \
--model mlx-community/Qwen3.6-27B-4bit \
--prompt "Explain mixture-of-experts in two sentences."
MLX takes a little more setup than Ollama, but if you are squeezing every token-per-second out of an M-series chip, it is the way to go. See our MLX framework guide for a deeper walkthrough.
That's it. You now have the highest-scoring Mac-runnable LLM running locally, tuned for your hardware. Qwen 3.6 handles coding, reasoning, and everyday tasks that used to require a cloud subscription — for free, and fully private.