

About
Continuous batching and SSD caching, managed from the menu bar.
From the repository: “LLM inference server with continuous batching & SSD caching for Apple Silicon — managed from the macOS menu bar”
Alternatives
All alternatives →OllamaRun open-weight LLMs locally with one commandDocker / self-hostDesktop90
JanOffline ChatGPT alternative that runs on your computerDesktop73
Stability MatrixPackage manager for Stable Diffusion UIsDesktop72
Text Generation WebUIDesktop app for running local LLMsWebDesktop69