By the llama.cpp team and Hugging Face

Your AI.
On your computer.

Llama is a tiny menu bar app that runs the latest open models on your Mac. Llama is a system tray app that runs the latest open models on your PC. Llama is an app for Mac and Windows that runs the latest open models on your computer. Chat with them, or use them in your other apps.

1 MB download · Free · Open source · Works offline

On Linux? Install the command-line version:

curl -LsSf https://llama.app/install.sh | sh

Prefer Brew or Winget? Package managers Rather build from source? Follow instructions

  1. 1 Open chat Chat with any model in your browser. Like ChatGPT, but on your Maccomputer.
  2. 2 A local API Your apps connect here. It speaks the OpenAI API, so coding agents, editors, and scripts just work.
  3. 3 Recommended for this Mac PC Models that fit your Maccomputer, one click to install. Llama picks the settings.

Like OpenAI, but on your Maccomputer

OpenAI has ChatGPT to chat with and an API to build on. Llama gives you both, running on your own computer, with models you choose.

For you

Chat · like ChatGPT

Click “Open chat” in the menu. Nothing to set up.

For your apps

API · like the OpenAI API

Coding agents Pi
Editors OpenAI-compatible
Chat apps OpenAI-compatible
Your own code Any OpenAI SDK

Point any app that works with OpenAI at localhost:9931/v1

Your models · download once, use everywhere

Qwen3.8 27B gpt-oss 20B gemma-4 E4B

Nothing to set up

Running AI locally used to mean reading forum threads about settings. Llama checks your Maccomputer and sets up each model to run well on it, based on how llama.cpp actually works. You only pick which model to talk to.

Know what you're doing? Every llama.cpp setting is still there, in one plain-text file.

Qwen3.8 27B Chosen by Llama, for your Maccomputer
  • Quantization Highest precision that fits in memory
  • Context size Only sizes that fit in memory
  • Image input On for models that understand images
  • Speculative decoding On for models that support it
  • Batch size Larger on Macs with 32 GB or more

Light enough to forget it's there

Local AI has a reputation for eating disk space and memory. Llama keeps both to a minimum.

4 MB app

A native Mac app, and just a 1 MB download — smaller than a photo.

Each model stored once

Kept in the Hugging Face cache, shared with llama.cpp and other tools.

Nothing loaded when idle

Models load when something asks for one and unload after 5 minutes idle.

A great model for every Maccomputer

Llama suggests one that fits when you open it. Here's where to start.

Browse models

Build on Llama

Your apps get a local AI API they can all share, and Llama takes care of the engine and the models.

OpenAI-compatible API

If your code works with OpenAI, it works with Llama. Change the base URL and keep everything else — no API keys, no usage bills.

  • OpenAI- and Anthropic-compatible endpoints
  • Streaming, tool calling, structured output, vision
  • Already use llama.cpp? Your models show up automatically
  • The llama command too: llama cli, llama serve, and more
  • Reach it from your other devices over Tailscale

$curl http://localhost:9931/v1/chat/completions \ -H "Content-Type: application/json" \ -d '{ "model": "ggml-org/gpt-oss-20b-GGUF:MXFP4", "messages": [{"role": "user", "content": "Hi!"}] }'

Nothing to bundle

Your app talks to Llama over the API, and a one-click link installs the model it needs. No engine to ship, no gigabytes in your download — and your users keep one copy of each model for all their apps.

Without Llama 57 GB on disk
Chat app
Engine Model 19 GB
Coding agent
Engine Model 19 GB
Your app
Engine Model 19 GB
With Llama 19 GB on disk
Chat appCoding agentYour app
Llama
Engine Model 19 GB

Local AI starts here

Free, open source, and yours to keep.

or
brew install --cask llama-app

Just want the command line?Install the command line:

curl -LsSf https://llama.app/install.sh | sh
irm https://llama.app/install.ps1 | iex