Performance
What is Time to First Token (TTFT)?
The latency between sending a prompt and receiving the first generated token. TTFT depends on prompt processing speed and prompt length. On Apple Silicon, TTFT for a 1,000-token prompt ranges from ~0.5 seconds (M5 Max, 8B model) to ~5 seconds (M3 Pro, 70B model). The M5 Max's Neural Accelerators significantly reduce TTFT for long-context prompts.
Where Time to First Token (TTFT) comes up on LLMCheck
- M5 Ultra and M6, Measured: How the LLMCheck Estimates Held Up (2026)
- M5 Max for Local AI: Complete Apple Silicon Benchmark Guide (2026)
- Phi-4 14B on M4: 12 tokens/sec — Benchmark & Setup (2026)
- Qwen 3 8B on M2: 17 tokens/sec — Benchmark & Setup (2026)
- LFM2.5-2.6B on M6: 79 tokens/sec — Benchmark & Setup (2026)
- Phi-4 14B on M6: 16 tokens/sec — Benchmark & Setup (2026)
Look up tok/s for your exact chip and model →