vLLM

vllm-project/vllm

High-throughput LLM serving engine

EnterpriseLocal inferenceDocker / self-hostLibrary / SDKPermissive
vllm.ai
Preview of vLLM
Preview of vLLM at vllm.ai

About

A fast inference and serving engine for LLMs built around PagedAttention and continuous batching. It is the standard choice for serving open models at scale on GPUs.

From the repository: “A high-throughput and memory-efficient inference and serving engine for LLMs”

SGLangFast serving framework for LLMsDocker / self-hostLibrary / SDK81
OllamaRun open-weight LLMs locally with one commandDocker / self-hostDesktop90
llama.cppLLM inference in C/C++ on CPUs and GPUsCLILibrary / SDK86
LocalAIDrop-in OpenAI API replacement that runs locallyDocker / self-host86

Topics