# AirLLM

> Run 70B models on a 4GB GPU

Layer-by-layer inference that lets large models run on small GPUs.

- Source: https://github.com/lyogavin/airllm
- License: Apache-2.0
- Language: Jupyter Notebook
- Stars: 35290
- Forks: 3726
- Contributors: 10
- Last commit: 2026-10-01
- Latest release: v4.0.0 (2026-09-05)
- Purpose: AI & LLM tooling, Local inference
- Runs on: Library / SDK
- For: Personal

## Worth score: 73/100

How much DigGitHub recommends it, from activity, adoption, docs, license and security signals.

- Popularity: 23/25
- Momentum: 10/20
- Maintenance: 25/25
- Community: 10/15
- Readiness: 5/15

## Alternatives

- [llama.cpp](https://diggithub.com/ggml-org/llama.cpp.md): LLM inference in C/C++ on CPUs and GPUs
- [vLLM](https://diggithub.com/vllm-project/vllm.md): High-throughput LLM serving engine
- [Unsloth](https://diggithub.com/unslothai/unsloth.md): Fast local fine-tuning and running of LLMs
- [SGLang](https://diggithub.com/sgl-project/sglang.md): Fast serving framework for LLMs
- [OpenVINO](https://diggithub.com/openvinotoolkit/openvino.md): Optimize and deploy AI inference
- [whisper.cpp](https://diggithub.com/ggml-org/whisper.cpp.md): Whisper speech recognition in C/C++
- [RunAnywhere SDKs](https://diggithub.com/RunanywhereAI/runanywhere-sdks.md): Run AI locally on devices
- [ncnn](https://diggithub.com/Tencent/ncnn.md): Neural network inference for mobile

---

Source page: https://diggithub.com/lyogavin/airllm
Updated: 2026-10-03
