Home/Glossary/Inference

Concept

What is Inference?

The process of running a trained AI model to generate text, code, or predictions. LLMCheck defines local inference as running this process entirely on your Mac's hardware without any server communication. Inference speed is measured in tokens per second and depends primarily on memory bandwidth and model size.

Where Inference comes up on LLMCheck

Browse all 81 models in the index →

Related terms

All 37 terms in the LLMCheck glossary →