
vLLM
vllm-project/vllmHigh-throughput LLM serving engine
EnterpriseLocal inferenceDocker / self-hostLibrary / SDKPermissive
About
A fast inference and serving engine for LLMs built around PagedAttention and continuous batching. It is the standard choice for serving open models at scale on GPUs.
From the repository: “A high-throughput and memory-efficient inference and serving engine for LLMs”