Performance
What is Time to First Token (TTFT)?
The latency between sending a prompt and receiving the first generated token. TTFT depends on prompt processing speed and prompt length. On Apple Silicon, TTFT for a 1,000-token prompt ranges from ~0.5 seconds (M5 Max, 8B model) to ~5 seconds (M3 Pro, 70B model). The M5 Max's Neural Accelerators significantly reduce TTFT for long-context prompts.
Where Time to First Token (TTFT) comes up on LLMCheck
- Phi-4 14B on M4: 12 tokens/sec — Benchmark & Setup (2026)
- Qwen 3 8B on M2: 17 tokens/sec — Benchmark & Setup (2026)
- LFM2.5-2.6B on M6: 81 tokens/sec — Benchmark & Setup (2026)
- Phi-4 14B on M6: 17 tokens/sec — Benchmark & Setup (2026)
- Qwen 3 8B on M1: 12 tokens/sec — Benchmark & Setup (2026)
- Qwen 3 14B on M3: 10 tokens/sec — Benchmark & Setup (2026)
Look up tok/s for your exact chip and model →