Prefill vs. Decode
Large language model inference has two distinct phases: prefill, when the model reads the prompt, and decode, when it produces the answer one token at a time. The same model runs in both phases, but the work changes enough that different hardware limits usually dominate.
Prefill: reading the prompt
During prefill, the model processes the input tokens in parallel. Each transformer layer computes representations for the prompt and stores the attention keys and values in the KV cache. The model then produces the probability distribution for the first output token.
So here’s how prefill looks like. We do one pass through the model: we read the weights once to perform the computations on GPU, generate the first token, and save the KV cache to memory for future use.
What affects the performance of prefill
Prefill is usually compute-bound, because the whole sequence is processed in parallel. “Compute-bound” is a fancy way to say that the more flops your GPU has, the faster it will do prefill. The following factors affect prefill performance:
- A longer prompt or longer context
- A larger or more computationally expensive model
- A slower processor or poorly optimized kernels
Decode: writing the answer
After choosing the first output token, the model enters an autoregressive loop. The model generates a token, then appends that token to the conversation, then it generates a new token. This process repeats until the end of the generation. For each generated token, the model weights and KV cache are read from memory to GPU, the token is generated and we write the new KV cache for the new token to memory.
Due to these repeated read/writes, decoding is memory-bandwidth-bound: the faster you can move stuff to the GPU, the faster the generation speed. There’s not much computation involved in generating a single token, so memory transfers dominate. Batching inputs, however, helps maximize compute. When we decode multiple generations together, we read the weights once for all the items in the batch, and the GPU generates the next tokens for all the conversations in parallel. This makes batched generation compute-bound, and we can get better throughput (the number of processed tokens increases).
Factors that affect decoding speed
- A larger model or higher-precision weights
- GPU memory bandwidth
- CPU/GPU offloading across a slow interconnect
- A long active context, which increases KV-cache reads
- An implementation without optimized kernels for the model and hardware
Decode speed is often reported as output tokens per second or time per output token.
KV Cache
KV cache is a trick to make attention run faster. Attention is the basic component of LLMs: each new token generation needs to look at all the previous tokens in the sequence. Therefore, it has quadratic complexity. Attention is composed of keys and values for every token, but once they have been calculated for a previous token in the sequence, we can cache the result and reuse it when generating a new token. This adds extra memory (which grows with context window) but makes inference much faster. In llama.cpp, the KV cache is pre-allocated for a given context window. In llama.app you can see the amount of memory a conversation will take including the model and the KV cache.
The memory consumption of KV cache depends on the model architecture, but we can approximate it with number of layers × context length × KV heads × head dim × precision.
Long conversations therefore have two costs:
- More memory is reserved or consumed by the KV cache.
- Each new token attends over more cached positions, increasing decode work and memory traffic.
Starting a fresh conversation, summarizing old turns, reducing the context limit, or quantizing the KV cache can help when supported by the runtime.
Metrics that map to user experience
Time to first token (TTFT)
From the moment you submit the initial prompt to the moment you see the first token, so it’s about prefill.
Time per output token (TPOT)
The time it takes between subsequent generated tokens, which gives a signal on how fast conversation feels.
Output throughput
The number of output tokens generated per second. You can benchmark this with llama bench.
# prefill 128 tokens, generate 64 tokens, run 3 times, benchmark throughput
llama bench -hf ggml-org/gemma-4-e4b-it-GGUF:Q4_0 -p 128 -n 64 -r 3 | model | size | params | backend | threads | test | t/s |
|---|---|---|---|---|---|---|
| gemma3 1B Q4_K | 762.49 MiB | 999.89 M | MTL,BLAS | 5 | pp128 | 2184.32 ± 3.54 |
| gemma3 1B Q4_K | 762.49 MiB | 999.89 M | MTL,BLAS | 5 | tg64 | 115.03 ± 0.19 |
The number of runs you pass (three, in this case) increases the accuracy of the estimate.
Aggregate throughput
The total tokens served across all active requests per second. Batching may improve this, even when it increases latency for an individual request.
End-to-end latency
The time from the moment you submit the prompt until the end of the generation, which is approximately TTFT + number of output tokens × TPOT.
Improving metrics
There are many tricks to improve throughput and memory, such as speculative decoding or quantization. They will be covered in additional guides.