

About
Packs a model and runtime into one executable that runs on most systems.
From the repository: “Distribute and run LLMs with a single file.”
Alternatives
All alternatives →OllamaRun open-weight LLMs locally with one commandDocker / self-hostDesktop90
llama.cppLLM inference in C/C++ on CPUs and GPUsCLILibrary / SDK86
whisper.cppWhisper speech recognition in C/C++CLILibrary / SDK79
ColibriRun large MoE models on everyday hardwareCLI79