Concept
What is Inference?
The process of running a trained AI model to generate text, code, or predictions. LLMCheck defines local inference as running this process entirely on your Mac's hardware without any server communication. Inference speed is measured in tokens per second and depends primarily on memory bandwidth and model size.
Where Inference comes up on LLMCheck
- Local AI Guides for Mac — Step-by-Step Setup & Installation
- Local AI Troubleshooting Hub — Fix Common LLM Issues on Mac
- How to Run Qwen 3.6 on a Mac (35B-A3B and 27B) — Setup Guide
- How to Install Ollama on Mac — Complete Setup Guide (2026)
- How to Fine-Tune a Local LLM on Mac with MLX (LoRA) — 2026 Guide
- How to Run Local LLMs on an Intel Mac (2026) — What's Possible & Realistic Speeds
Browse all 81 models in the index →