Metric
What is Tokens Per Second (tok/s)?
The primary speed metric for LLM inference, measuring how many tokens the model generates per second. According to the LLMCheck index on M5 Max 128 GB: Phi-4 Mini ~135 tok/s, Qwen 3.5 9B ~100 tok/s, Qwen 3 30B-A3B ~58 tok/s, Llama 4 Scout ~32 tok/s, DeepSeek R1 70B ~10 tok/s. Above 30 tok/s feels like real-time conversation.
Where Tokens Per Second (tok/s) comes up on LLMCheck
- Local AI Guides for Mac — Step-by-Step Setup & Installation
- How to Build a Local RAG System on Mac with Ollama
- Local AI Troubleshooting Hub — Fix Common LLM Issues on Mac
- How to Run Qwen 3.6 on a Mac (35B-A3B and 27B) — Setup Guide
- How to Use MCP (Model Context Protocol) with Local LLMs on Mac (2026)
- How to Install Ollama on Mac — Complete Setup Guide (2026)
Look up tok/s for your exact chip and model →