Your AI.
On your computer.
Llama is a tiny menu bar app that runs the latest open models on your Mac. Llama is a system tray app that runs the latest open models on your PC. Llama is an app for Mac and Windows that runs the latest open models on your computer. Chat with them, or use them in your other apps.
1 MB download · Free · Open source · Works offline
On Linux? Install the command-line version:
curl -LsSf https://llama.app/install.sh | sh Prefer Brew or Winget? Package managers Rather build from source? Follow instructions
- 1 Open chat Chat with any model in your browser. Like ChatGPT, but on your Maccomputer.
- 2 A local API Your apps connect here. It speaks the OpenAI API, so coding agents, editors, and scripts just work.
- 3 Recommended for this Mac PC Models that fit your Maccomputer, one click to install. Llama picks the settings.
Like OpenAI, but on your Maccomputer
OpenAI has ChatGPT to chat with and an API to build on. Llama gives you both, running on your own computer, with models you choose.
For you
Chat · like ChatGPT
Click “Open chat” in the menu. Nothing to set up.
For your apps
API · like the OpenAI API
Point any app that works with OpenAI at localhost:9931/v1
Your models · download once, use everywhere
Nothing to set up
Running AI locally used to mean reading forum threads about settings. Llama checks your Maccomputer and sets up each model to run well on it, based on how llama.cpp actually works. You only pick which model to talk to.
Know what you're doing? Every llama.cpp setting is still there, in one plain-text file.
- Quantization Highest precision that fits in memory
- Context size Only sizes that fit in memory
- Image input On for models that understand images
- Speculative decoding On for models that support it
- Batch size Larger on Macs with 32 GB or more
Light enough to forget it's there
Local AI has a reputation for eating disk space and memory. Llama keeps both to a minimum.
4 MB app
A native Mac app, and just a 1 MB download — smaller than a photo.
Each model stored once
Kept in the Hugging Face cache, shared with llama.cpp and other tools.
Nothing loaded when idle
Models load when something asks for one and unload after 5 minutes idle.
A great model for every Maccomputer
Llama suggests one that fits when you open it. Here's where to start.
Build on Llama
Your apps get a local AI API they can all share, and Llama takes care of the engine and the models.
OpenAI-compatible API
If your code works with OpenAI, it works with Llama. Change the base URL and keep everything else — no API keys, no usage bills.
- OpenAI- and Anthropic-compatible endpoints
- Streaming, tool calling, structured output, vision
- Already use llama.cpp? Your models show up automatically
- The
llamacommand too:llama cli,llama serve, and more - Reach it from your other devices over Tailscale
$curl http://localhost:9931/v1/chat/completions \ -H "Content-Type: application/json" \ -d '{ "model": "ggml-org/gpt-oss-20b-GGUF:MXFP4", "messages": [{"role": "user", "content": "Hi!"}] }'
Nothing to bundle
Your app talks to Llama over the API, and a one-click link installs the model it needs. No engine to ship, no gigabytes in your download — and your users keep one copy of each model for all their apps.
Local AI starts here
Free, open source, and yours to keep.
brew install --cask llama-app Just want the command line?Install the command line:
curl -LsSf https://llama.app/install.sh | sh irm https://llama.app/install.ps1 | iex