
llama.cpp
ggml-org/llama.cppLLM inference in C/C++ on CPUs and GPUs
PersonalLocal inferenceCLILibrary / SDKPermissive
About
The inference engine behind much of the local-LLM ecosystem. It runs quantized GGUF models efficiently on laptops, servers and phones, and ships an OpenAI-compatible server.
From the repository: “LLM inference in C/C++”
Alternatives
All alternatives →whisper.cppWhisper speech recognition in C/C++CLILibrary / SDK79
OllamaRun open-weight LLMs locally with one commandDocker / self-hostDesktop90
vLLMHigh-throughput LLM serving engineDocker / self-hostLibrary / SDK86
UnslothFast local fine-tuning and running of LLMsWebLibrary / SDK81