Colibri

JustVugg/colibri

Run large MoE models on everyday hardware

PersonalLocal inferenceCLIPermissive
justvugg.github.io
Preview of Colibri
Preview of Colibri at justvugg.github.io

About

A pure C engine that streams expert weights from disk to run big mixture-of-experts models on modest machines.

From the repository: “Run frontier MoE models on hardware you already own — pure C, zero deps, experts streamed from disk. Tiny engine, immense model. 🐦”

OllamaRun open-weight LLMs locally with one commandDocker / self-hostDesktop90
llama.cppLLM inference in C/C++ on CPUs and GPUsCLILibrary / SDK86
whisper.cppWhisper speech recognition in C/C++CLILibrary / SDK79
llmfitFind which LLMs run on your hardwareCLI79