AirLLM

lyogavin/airllm

Run 70B models on a 4GB GPU

PersonalLocal inferenceLibrary / SDKPermissive
github.com
Preview of AirLLM
Preview of AirLLM

About

Layer-by-layer inference that lets large models run on small GPUs.

From the repository: “AirLLM 70B inference with single 4GB GPU”

llama.cppLLM inference in C/C++ on CPUs and GPUsCLILibrary / SDK86
vLLMHigh-throughput LLM serving engineDocker / self-hostLibrary / SDK86
UnslothFast local fine-tuning and running of LLMsWebLibrary / SDK81
SGLangFast serving framework for LLMsDocker / self-hostLibrary / SDK81

Topics