Home/Glossary/llama.cpp

Software

What is llama.cpp?

An open-source C/C++ library for running LLM inference on consumer hardware. llama.cpp is the inference engine behind Ollama and many other Mac AI apps. It supports Apple Silicon Metal GPU acceleration and GGUF model files. According to the LLMCheck index, llama.cpp delivers baseline performance that Apple's MLX framework exceeds by 20–50% on Apple Silicon.

Where llama.cpp comes up on LLMCheck

Compare the 13 apps that run LLMs on a Mac →

Related terms

All 37 terms in the LLMCheck glossary →