The stack at a glance
A local coding assistant has two pieces: a model server that runs the LLM on your hardware, and an editor integration that feeds it your code and shows you suggestions. We use Ollama for the server and Qwen3.6-27B for the model — Alibaba's Apache 2.0 model that, according to the LLMCheck index, hits 77.2% on SWE-bench Verified, the best of any open model you can run on a 24 GB Mac.[1]
It is a dense 27B model — every parameter works on every token, which is what buys those benchmark scores. At Q4 it fits in ~16 GB and runs at ~40 tok/s (estimated) on an M5 Max. If you want faster generation on the same RAM, its MoE sibling Qwen3.6-35B-A3B (only ~3B parameters active per token, 73.4% SWE-bench) is the quick-draw alternative. For the editor, you have three good choices, and this guide covers all of them: Continue.dev (VS Code), Zed (native Mac editor with built-in AI), and Cursor (pointed at a custom endpoint).
| Editor | Best for | Offline? |
|---|---|---|
| Continue.dev | Deepest local-model control in VS Code | Fully local |
| Zed | Fast native editor, least setup | Fully local |
| Cursor | Already a Cursor user | Mostly (some cloud features) |
Step 1: Hardware Check
Qwen3.6-27B needs about 16 GB of unified memory at Q4, plus headroom for your editor and the code context you feed it. That makes a 24 GB Mac the recommended minimum. Check yours under Apple menu → About This Mac → Memory.
- 24 GB+ — ideal. Run the full 27B model.
- 16 GB — drop to a compact coder such as
nanbeige4.2:3b(63.6% SWE-bench Verified, vendor-reported), which fits comfortably and is still strong for autocomplete and routine edits. - 32 GB+ — raise the context window for whole-file and multi-file work.
Buying a Mac for local coding? A 24–32 GB M4 Pro or M5 is the value sweet spot. Our Mac hardware buying hub ranks configurations by real-world tok/s per dollar so you don't overspend on memory you won't use — or under-buy and stall.
Step 2: Install Ollama & Pull Qwen 3.6
Install Ollama from ollama.com (full walkthrough in our Install Ollama on Mac guide), then pull the model:
ollama pull qwen3.6:27b
This downloads the Q4 build (~16 GB). To confirm it is ready and check that Ollama's local server is serving it:
ollama list # should show qwen3.6:27b
ollama serve # starts the API at http://localhost:11434 (usually already running)
Ollama normally runs its server automatically in the background, so ollama serve is only needed if the API isn't already up. You can do a quick smoke test:
curl http://localhost:11434/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{"model":"qwen3.6:27b","messages":[{"role":"user","content":"Write a Python one-liner to flatten a nested list."}]}'
Step 3: Install Continue.dev (or Zed / Cursor)
Pick the editor that fits how you work. The rest of this guide uses Continue.dev as the primary path because it offers the most control, with Zed and Cursor notes alongside.
Continue.dev (VS Code)
Open VS Code, go to the Extensions panel (Cmd+Shift+X), search for Continue, and click Install. A Continue icon appears in your sidebar. It also works in JetBrains IDEs via their plugin marketplace.
Zed
Download Zed from zed.dev. AI is built in — no extension needed. You'll configure its Ollama provider in the next step.
Cursor
If you already use Cursor, you can keep it and point it at your local model. Open Settings → Models, enable a custom OpenAI-compatible base URL, and you'll wire it up in Step 4.
Step 4: Point the Extension at Your Local Model
Continue.dev config
Open the Continue config (click the gear in the Continue sidebar, or edit ~/.continue/config.yaml) and add Qwen 3.6 as both your chat model and your autocomplete model:
models:
- name: Qwen 3.6 (local)
provider: ollama
model: qwen3.6:27b
apiBase: http://localhost:11434
roles:
- chat
- edit
- apply
- name: Qwen 3.6 Autocomplete
provider: ollama
model: qwen3.6:27b
apiBase: http://localhost:11434
roles:
- autocomplete
Save the file. Continue picks up the change immediately — you'll see "Qwen 3.6 (local)" in the model dropdown.
Zed config
In Zed, open settings.json (Cmd+,) and register Ollama as a language-model provider:
{
"language_models": {
"ollama": {
"api_url": "http://localhost:11434"
}
},
"assistant": {
"default_model": {
"provider": "ollama",
"model": "qwen3.6:27b"
}
}
}
Cursor config
In Cursor's Settings → Models, add a custom model with the OpenAI-compatible base URL pointing at Ollama:
Base URL: http://localhost:11434/v1
API Key: ollama (any non-empty string works locally)
Model: qwen3.6:27b
Heads-up on Cursor: some Cursor features (like its tab autocomplete and indexing) still route through Cursor's cloud even with a custom model. For a guaranteed fully-offline assistant, prefer Continue.dev or Zed.
Step 5: Use It — Autocomplete, Chat & Agent Mode
With the model wired in, you now have a full coding assistant running on-device. Three things to try:
- Tab autocomplete — start typing a function and Qwen 3.6 suggests the rest inline. Press
Tabto accept. Great for boilerplate, tests, and repetitive patterns. - Inline chat & edit — select code and press
Cmd+I(Continue) to ask for a refactor, a bug fix, or an explanation. The model rewrites the selection in place and shows a diff you can accept or reject. - Agent mode — in Continue's chat, switch to Agent and give a higher-level task ("add input validation to this endpoint and a test for it"). It reads relevant files, proposes multi-file changes, and applies them with your approval.
A good first prompt to feel it out, with a file open:
Refactor this function to use early returns and add a docstring.
Then write a pytest test that covers the edge cases.
Everything here runs through your local Ollama server — no network calls, no data sent to any vendor, and it keeps working on a plane or behind a firewall.
Tip: For autocomplete that stays snappy, some teams pair a small fast model (e.g. nanbeige4.2:3b) for tab completion with the 27B model for chat and agent mode. Continue.dev lets you assign different models per role, exactly as shown in Step 4.
Step 6: Tips & Scaling Beyond Your Mac
Give it more context on bigger Macs
The more of your codebase the model can see, the better its edits. On 32 GB+ Macs, raise the context window so it can hold larger files and more surrounding code. With Ollama you can bake a larger window into a custom model:
# Modelfile
FROM qwen3.6:27b
PARAMETER num_ctx 32768
ollama create qwen3.6-32k -f Modelfile
# then reference qwen3.6-32k in your editor config
On a 24 GB Mac, keep context moderate (8K–16K) so the model doesn't swap and slow down.
When you need a coder bigger than your Mac
Some 2026 frontier coders — DeepSeek V4 Pro, Kimi K3 — are simply too large for any Mac's unified memory. When a hard task outruns what Qwen3.6-27B can do locally, the practical option is to rent a GPU by the hour and run the big model there, keeping the same Ollama/OpenAI-compatible workflow — just point your editor at the remote endpoint instead of localhost.
A cost-effective place to do this is Vast.ai, a marketplace for on-demand GPUs where an H100 or 80GB card runs a few dollars an hour — far cheaper than buying hardware for an occasional frontier-model task.
Disclosure: the Vast.ai link is a referral; if you sign up through it, LLMCheck may earn a small credit at no extra cost to you. We only recommend it because renting beats buying for occasional big-model jobs.
You're set. You now have a private coding assistant — autocomplete, chat, and agent edits — running entirely on your Mac, with a clear path to rent extra horsepower only when you actually need it.