Prompt caching in vLLM reuses the work already done on a shared prompt prefix.
On one NVIDIA RTX PRO 6000, a 32B model with a 32,000-token shared prompt went from 7.8 seconds to 0.15 seconds to first token and used 68% less energy per request.
The saving grows with the shared prefix: 21% at 2,000 tokens, 82% at 32,000 on a 7B.
LeanLM (not affiliated with Google's LearnLM educational AI) measures what AI workloads cost to serve. Every number below comes from our own runs on a stated GPU.
What does prefix caching do?
Prefix caching skips recomputing a prompt the server has already seen. With it on, the server keeps the KV cache (the attention keys and values the model computed for those tokens) for a prompt prefix it has already processed and reuses it when the next request starts with the same tokens, so it skips prefill (the first pass over the prompt tokens) on those tokens. That's why the saving grows with the length of the shared prefix and why unique prompts don't benefit. Output tokens still cost the same regardless, which is why the 32B at 2,000 tokens, where most of each request's time goes to writing the answer, only saved 11%.
Prompt caching vs prefix caching
Prompt caching and prefix caching are the same mechanism run by different operators. Provider "prompt caching" from Anthropic, OpenAI and Google is the idea run by the API provider and sold as a discount on cached input tokens. "Prefix caching" is what vLLM calls it when you run the server yourself, with --enable-prefix-caching instead of a billing line. See provider prompt caching pricing for what the big three charge for it. So vLLM prompt caching and vLLM prefix caching are the same feature under two names.
How much does prompt caching save in vLLM?
vLLM prefix caching performance, measured: it cut energy per request 11% to 82% and time to first token (TTFT) by up to 53 times, depending on how long the shared prompt is. We ran Qwen2.5 7B and Qwen2.5 32B on vLLM 0.30.0 with prefix caching off and on, at three shared-prompt lengths, one request at a time:
| Model | Shared prompt | TTFT off / on | Latency off / on | J per request off / on | Energy change | req/s off / on |
|---|---|---|---|---|---|---|
| 7B | 2k tokens | 0.090 s / 0.025 s | 0.71 s / 0.65 s | 251 J / 199 J | -21% | 1.20 / 1.39 |
| 7B | 8k tokens | 0.353 s / 0.033 s | 0.92 s / 0.62 s | 365 J / 184 J | -49% | 0.99 / 1.47 |
| 7B | 32k tokens | 1.937 s / 0.083 s | 2.42 s / 0.59 s | 1,188 J / 212 J | -82% | 0.38 / 1.23 |
| 32B | 2k tokens | 0.411 s / 0.100 s | 4.94 s / 4.73 s | 1,600 J / 1,429 J | -11% | 0.20 / 0.22 |
| 32B | 8k tokens | 1.559 s / 0.090 s | 7.56 s / 6.20 s | 2,209 J / 1,426 J | -35% | 0.16 / 0.21 |
| 32B | 32k tokens | 7.819 s / 0.148 s | 14.46 s / 6.79 s | 5,909 J / 1,906 J | -68% | 0.07 / 0.16 |
TTFT (time to first token) and latency are medians; J per request is idle-subtracted energy (idle power 84 to 92 W with the model loaded, measured for 60 s and subtracted). One NVIDIA RTX PRO 6000 Blackwell Server Edition (96 GB), vLLM 0.30.0, one request at a time. Full data: prompt-caching-vllm-ollama-measured.csv.
The rig runs in your VPC, or your engineers run it and we interpret the output.
Does batching stack with caching?
Batching and caching stack, and together they cut energy per request the most. At an 8,000-token prompt with 8 requests served at once, the 7B model used 206 J per request with the cache off (versus 365 J one at a time, -44%) at a median latency of 3.08 s, and 30 J with the cache on (-92% versus one at a time with the cache off) at 0.69 s latency and 9.18 req/s. The 32B used 999 J with the cache off (versus 2,209 J, -55%) at 20.5 s latency, and 194 J with the cache on (-91%) at 5.98 s latency and 1.55 req/s.
Does Ollama cache prompts?
Yes, Ollama reuses the prefix from the previous request in the same slot by default. On qwen2.5:7b-instruct at its default Q4_K_M with an 8,000-token prompt, one request at a time, the default reuse gave a TTFT of 0.19 s, latency of 0.53 s and 185 J raw per request. Defeating that reuse, by changing the first token of the system prompt on every request, pushed TTFT to 1.03 s, latency to 1.34 s and raw energy to 488 J. Keeping the default reuse cut raw energy per request 62%. We report raw (not idle-subtracted) energy for Ollama because the two arms measured different idle power, 34 W versus 87 W.
How do you enable prefix caching in vLLM?
To enable prompt caching in vLLM, pass --enable-prefix-caching on the serve command:
vllm serve Qwen/Qwen2.5-7B-Instruct --enable-prefix-caching
Recent vLLM versions may turn this on by default, so check your version's docs at docs.vllm.ai before assuming it's off. Either way, keep shared content at the start of the prompt (system prompt, tool definitions, long documents) and per-request content at the end (the user's question, timestamps, request IDs). A match has to be token for token, so anything that moves earlier in the prompt breaks it.
When doesn't caching help?
Caching doesn't help on short or unique prompts, since there's little or no shared prefix to skip. It also helps less when long outputs dominate the request, because output tokens are generated one at a time regardless of what's cached and most of each request's time goes to writing the answer, which is why the 32B at a 2,000-token prompt only saved 11%. And it never helps on the first request with a new prefix, since that one still pays full price to build the cache entry.
Setup problems we hit
Four things slowed us down running this test, in case they slow you down too:
- vLLM 0.30.0 crashed at engine start on this Blackwell card with "FlashInfer requires GPUs with sm75 or higher." Setting
VLLM_USE_FLASHINFER_SAMPLER=0andVLLM_ATTENTION_BACKEND=FLASH_ATTNfixed it. - Qwen2.5 refuses
--max-model-lenabove 32768, its config limit, unless you override it. - The first request with a new prefix pays full price. Measure after a warm-up request and exclude it.
- Ollama reuses the prefix from the previous request in the same slot by default. Changing anything at the start of the prompt (a timestamp or a request ID) defeats it.
What hasn't been measured yet?
This test covers one GPU, two model sizes and concurrency up to 8. It doesn't cover production concurrency beyond that, real RAG corpora instead of a single shared document, AMD GPUs, SGLang, cache eviction under memory pressure, or multi-GPU serving.
Frequently Asked Questions
Does vLLM support prompt caching?
Yes. vLLM calls it automatic prefix caching, turned on with --enable-prefix-caching. On a 32,000-token shared prompt it cut time to first token from 7.8 s to 0.15 s on a 32B model in our test.
What is the difference between prompt caching and prefix caching?
Same idea, different operator. Providers like Anthropic and OpenAI call it prompt caching and bill cached input tokens at a discount. vLLM calls it prefix caching when you run the server yourself.
How much energy does prompt caching save?
It depends on how long the shared prompt is. On a 7B model it saved 21% of energy per request with a 2,000-token prefix and 82% with a 32,000-token prefix, measured on one NVIDIA RTX PRO 6000.
Does Ollama cache prompts?
Yes, by default it reuses the prefix from the previous request. Keeping that reuse cut energy per request 62% and time to first token from 1.03 s to 0.19 s on a 7B model with an 8,000-token prompt.
Why is prefix caching not working?
The prefix has to match token for token. Anything that changes at the start of the prompt, like a timestamp or request ID, breaks the match. Put shared content first and per-request content last.
Does batching reduce energy per request?
Yes. Serving 8 requests at once instead of 1 cut energy per request 44% on a 7B model and 55% on a 32B. With prefix caching as well, the 7B went from 365 J to 30 J per request.
Methods
- Hardware. RunPod, one NVIDIA RTX PRO 6000 Blackwell Server Edition (96 GB), 2026-09-28.
- Software. vLLM 0.30.0, with
VLLM_USE_FLASHINFER_SAMPLER=0andVLLM_ATTENTION_BACKEND=FLASH_ATTNset to work around a FlashInfer crash on this card. - Models. Qwen/Qwen2.5-7B-Instruct and Qwen/Qwen2.5-32B-Instruct in BF16, max_model_len 32768, gpu_memory_utilization 0.90.
- Workload. A shared prefix of about 2,000, 8,000 or 32,000 tokens (a system prompt plus a public-domain book, Pride and Prejudice from Project Gutenberg) followed by 100 different short questions; temperature 0, max_tokens 128.
- Warm-up. One warm-up request per arm, excluded from the measurement, since it's what fills the cache in the cache-on arms.
- Cache setting. Cache on:
--enable-prefix-caching; cache off:--no-enable-prefix-caching; the server was restarted between arms. - Energy method. nvidia-smi power.draw sampled at 5 Hz, GPU board power. Idle power (84 to 92 W with the model loaded) measured for 60 s and subtracted. Energy per request is energy over the arm divided by requests.
- Ollama method. qwen2.5:7b-instruct at its default Q4_K_M, 8,000-token prompt, one request at a time. Energy is raw (not idle-subtracted) because the two arms measured different idle power, 34 W versus 87 W.
- Runs. Single run per arm, not averaged across repeats.
- Data. Every measured row on this page: prompt-caching-vllm-ollama-measured.csv.
- Date. Measured 2026-09-28.