Llama 0.6.0 adds Clef and Qwen support, dramatic Metal speedup on Apple Silicon
The new version of Llama, the menu-bar app that runs open models locally on Mac and Windows, lands with four core updates: support for the Clef text-and-vision model, quality-tuned support for Qwen3.8-Flash-Next, a large performance jump on Metal for Apple Silicon, and a new batch API called `llama_batch_ext`. The site has been redesigned around the tagline "your intelligence, on your computer" — 1 MB download, open source, works offline.
Llama sits in the menu bar on macOS or the system tray on Windows and exposes two interfaces at once: a local ChatGPT-style chat that opens on click, and an OpenAI-compatible API running on `localhost:9931/v1`. Any application that speaks the OpenAI protocol — code editors, coding agents, scripts — points at that address and works without keys, accounts, or cloud. Models download once, live in the Hugging Face cache, and are shared by every tool that uses llama.cpp.
The engine upgrade in 0.6.0 is Clef support, a multimodal model that understands text and images, plus dedicated quality handling for the new Qwen3.8-Flash-Next. On Macs with Apple Silicon the developers report a "massive performance improvement" thanks to Metal optimizations, and the new `llama_batch_ext` API enables more efficient batch processing for heavy workloads. Every llama.cpp setting remains available in a plain-text config for anyone who wants to control quantization, context size, speculative decoding, and batch size — which now grows automatically on Macs with 32 GB of RAM or more.
The app weighs 4 MB on Mac (1 MB download), loads nothing at idle, and unloads the model after five minutes of inactivity. On first launch it scans the hardware and suggests a model that fits: maximum precision that fits in RAM, maximum context that fits, and image support if the model provides it. Models already downloaded via llama.cpp or other tools appear automatically, with no disk duplication.
For developers, the exposed endpoint is compatible with both OpenAI and Anthropic, including streaming, tool calls, structured output, and vision. The CLI ships with `llama cli`, `llama serve`, and more, and the API can be reached from other devices over Tailscale. Applications built on Llama don't need to bundle an engine or gigabytes of weights; they talk to the API, and a single install link pulls the required model for the end user. The user keeps one copy for every application.