Home/Glossary/Prompt Processing (Prefill)

Performance

What is Prompt Processing (Prefill)?

The initial phase where the model reads and processes your entire input prompt before generating the first token. Prompt processing speed is measured in tokens per second and depends heavily on GPU compute power. The M5 Max's 4x peak AI compute over M4 Max makes it significantly faster at processing long prompts, especially for coding and RAG workflows with large context.

Where Prompt Processing (Prefill) comes up on LLMCheck

Look up tok/s for your exact chip and model →

Related terms

All 37 terms in the LLMCheck glossary →