
Colibri
JustVugg/colibriRun large MoE models on everyday hardware
PersonalLocal inferenceCLIPermissiveAbout
A pure C engine that streams expert weights from disk to run big mixture-of-experts models on modest machines.
From the repository: “Run frontier MoE models on hardware you already own — pure C, zero deps, experts streamed from disk. Tiny engine, immense model. 🐦”
Alternatives
All alternatives →OllamaRun open-weight LLMs locally with one commandDocker / self-hostDesktop90
llama.cppLLM inference in C/C++ on CPUs and GPUsCLILibrary / SDK86
whisper.cppWhisper speech recognition in C/C++CLILibrary / SDK79
llmfitFind which LLMs run on your hardwareCLI79