llama.cpp

ggml-org/llama.cpp

LLM inference in C/C++ on CPUs and GPUs

PersonalLocal inferenceCLILibrary / SDKPermissive
llama.app
Preview of llama.cpp
Preview of llama.cpp at llama.app

About

The inference engine behind much of the local-LLM ecosystem. It runs quantized GGUF models efficiently on laptops, servers and phones, and ships an OpenAI-compatible server.

From the repository: “LLM inference in C/C++”

whisper.cppWhisper speech recognition in C/C++CLILibrary / SDK79
OllamaRun open-weight LLMs locally with one commandDocker / self-hostDesktop90
vLLMHigh-throughput LLM serving engineDocker / self-hostLibrary / SDK86
UnslothFast local fine-tuning and running of LLMsWebLibrary / SDK81

Topics