# vLLM

> High-throughput LLM serving engine

A fast inference and serving engine for LLMs built around PagedAttention and continuous batching. It is the standard choice for serving open models at scale on GPUs.

- Source: https://github.com/vllm-project/vllm
- Homepage: https://vllm.ai
- Docs: https://docs.vllm.ai
- License: Apache-2.0
- Language: Python
- Stars: 93132
- Forks: 22967
- Contributors: 3557
- Last commit: 2026-10-04
- Latest release: v0.30.0 (2026-09-22)
- Purpose: AI & LLM tooling, Local inference
- Runs on: Docker / self-host, Library / SDK
- For: Enterprise

## Worth score: 86/100

How much DigGitHub recommends it, from activity, adoption, docs, license and security signals.

- Popularity: 25/25
- Momentum: 10/20
- Maintenance: 25/25
- Community: 15/15
- Readiness: 11/15

## Alternatives

- [SGLang](https://diggithub.com/sgl-project/sglang.md): Fast serving framework for LLMs
- [Ollama](https://diggithub.com/ollama/ollama.md): Run open-weight LLMs locally with one command
- [Unsloth](https://diggithub.com/unslothai/unsloth.md): Fast local fine-tuning and running of LLMs
- [ncnn](https://diggithub.com/Tencent/ncnn.md): Neural network inference for mobile
- [OpenVINO](https://diggithub.com/openvinotoolkit/openvino.md): Optimize and deploy AI inference
- [RunAnywhere SDKs](https://diggithub.com/RunanywhereAI/runanywhere-sdks.md): Run AI locally on devices
- [LocalAI](https://diggithub.com/mudler/LocalAI.md): Drop-in OpenAI API replacement that runs locally
- [AirLLM](https://diggithub.com/lyogavin/airllm.md): Run 70B models on a 4GB GPU

---

Source page: https://diggithub.com/vllm-project/vllm
Updated: 2026-10-04
