# llama.cpp

> LLM inference in C/C++ on CPUs and GPUs

The inference engine behind much of the local-LLM ecosystem. It runs quantized GGUF models efficiently on laptops, servers and phones, and ships an OpenAI-compatible server.

- Source: https://github.com/ggml-org/llama.cpp
- Homepage: https://llama.app
- License: MIT
- Language: C++
- Stars: 130238
- Forks: 24051
- Contributors: 2049
- Last commit: 2026-10-04
- Latest release: v0.5.0 (2026-09-23)
- Purpose: AI & LLM tooling, Local inference
- Runs on: CLI, Library / SDK
- For: Personal

## Worth score: 86/100

How much DigGitHub recommends it, from activity, adoption, docs, license and security signals.

- Popularity: 25/25
- Momentum: 10/20
- Maintenance: 25/25
- Community: 15/15
- Readiness: 11/15

## Alternatives

- [Ollama](https://diggithub.com/ollama/ollama.md): Run open-weight LLMs locally with one command
- [vLLM](https://diggithub.com/vllm-project/vllm.md): High-throughput LLM serving engine
- [Unsloth](https://diggithub.com/unslothai/unsloth.md): Fast local fine-tuning and running of LLMs
- [SGLang](https://diggithub.com/sgl-project/sglang.md): Fast serving framework for LLMs
- [OpenVINO](https://diggithub.com/openvinotoolkit/openvino.md): Optimize and deploy AI inference
- [Colibri](https://diggithub.com/JustVugg/colibri.md): Run large MoE models on everyday hardware
- [llmfit](https://diggithub.com/AlexsJones/llmfit.md): Find which LLMs run on your hardware
- [AirLLM](https://diggithub.com/lyogavin/airllm.md): Run 70B models on a 4GB GPU

---

Source page: https://diggithub.com/ggml-org/llama.cpp
Updated: 2026-10-04
