llamafile

mozilla-ai/llamafile

Run an LLM from a single file

PersonalLocal inferenceCLIOther / custom
docs.mozilla.ai
Preview of llamafile
Preview of llamafile at docs.mozilla.ai

About

Packs a model and runtime into one executable that runs on most systems.

From the repository: “Distribute and run LLMs with a single file.”

OllamaRun open-weight LLMs locally with one commandDocker / self-hostDesktop90
llama.cppLLM inference in C/C++ on CPUs and GPUsCLILibrary / SDK86
whisper.cppWhisper speech recognition in C/C++CLILibrary / SDK79
ColibriRun large MoE models on everyday hardwareCLI79

Topics