Home/Glossary/Tokens Per Second (tok/s)

Metric

What is Tokens Per Second (tok/s)?

The primary speed metric for LLM inference, measuring how many tokens the model generates per second. According to the LLMCheck index on M5 Max 128 GB: Phi-4 Mini ~135 tok/s, Qwen 3.5 9B ~100 tok/s, Qwen 3 30B-A3B ~58 tok/s, Llama 4 Scout ~32 tok/s, DeepSeek R1 70B ~10 tok/s. Above 30 tok/s feels like real-time conversation.

Where Tokens Per Second (tok/s) comes up on LLMCheck

Look up tok/s for your exact chip and model →

Related terms

All 37 terms in the LLMCheck glossary →