Prompt Caching Explained: Cutting Inference Costs Without Cutting Quality
How prompt caching actually works at the API level, when it saves real money, and the setup mistakes that silently disable it.
Jump to 6 sections
Quick answer: Prompt caching stores the processed representation of a repeated prompt prefix — like a long system prompt or document — so a model provider doesn't re-run the expensive part of inference on every call. It typically cuts cost on the cached portion by 50-90% and cuts latency too, but it only helps when your prompts actually share a stable prefix.
This review covers how prompt caching actually works at the API level, when it saves real money versus when it's a rounding error, and the mistakes that quietly disable it without you noticing. It's for anyone shipping an LLM feature with a system prompt, tool schema, or reference document that repeats across calls.
How Prompt Caching Actually Works
Every call to a large language model starts by processing the input tokens before generating a single output token — this is the "prefill" step, and it scales roughly with prompt length. Prompt caching lets a provider skip re-computing that prefill for tokens it has already processed recently, as long as those tokens appear in the exact same order at the start of the prompt.
At The Model Drop, we've benchmarked cached versus uncached calls across several major providers, and the pattern is consistent: a cache hit on a long system prompt or reference document can cut time-to-first-token by half or more, on top of the direct cost savings on cached input tokens.
The catch is the word "prefix." Caching keys off a shared starting sequence of tokens, not the prompt as a whole. If anything before your reusable content changes between calls — a timestamp, a session ID, a reordered instruction — the cache misses entirely for that call, and you pay full price with no warning in most APIs.
When Caching Actually Saves Money
Caching pays off when three conditions line up: a long, stable prefix; enough call volume to reuse it within the cache's lifetime; and variable content placed after the cached portion, not mixed into it.
A worked example makes this concrete. Say a support agent sends a 3,000-token system prompt plus tool schema on every call, with only the user's message changing. Without caching, that 3,000-token prefix is billed at full input-token price every single call. With caching, only the first call in a given cache window pays full price; every call after that inside the cache lifetime pays the discounted cached rate on those same 3,000 tokens.
Where caching stops mattering is low-volume, high-variance workloads — a research assistant where every prompt is genuinely different, or an endpoint that gets one call every few minutes, well outside typical cache lifetimes (usually 5-10 minutes across providers, though this varies and changes over time). In our own cost audits, we've seen teams add caching to a low-traffic internal tool and see almost no bill impact, because the cache had already expired by the time the next call came in.
The underlying mechanics trace back to how transformer inference splits into a prefill phase and a decode phase — a distinction covered in the original attention-mechanism research indexed on arXiv. Prefill is the parallelizable, cache-friendly part; decode, which generates tokens one at a time, is not something caching speeds up.
The Setup Mistakes That Silently Disable It
Caching failures are usually invisible unless you're checking cache-hit metrics directly, since an uncached call still returns a correct answer — it's just full price.
- Variable content before the cacheable prefix. A timestamp, request ID, or user name placed at the very top of the prompt breaks the match for everything after it. Move stable content first.
- Non-deterministic prompt construction. If your prompt-building code reorders a list of tools or documents differently on each call (e.g. from an unsorted dictionary), the token sequence changes even though the content is logically the same.
- Cache lifetime mismatches. Sending cacheable calls further apart than the provider's cache TTL means every call is effectively a fresh cache write, not a hit.
- Not reading provider-reported cache token counts. Most APIs return how many tokens were served from cache per response — ignoring that field means shipping a caching setup you can't verify is actually working.
IEEE Spectrum has reported on how quickly inference costs can scale for teams that don't audit these details, since a broken cache doesn't throw an error — it just quietly bills every call at the uncached rate until someone checks the dashboard.
Caching vs. Other Cost Levers
Prompt caching is one of several cost levers, and it's worth knowing where it fits relative to the others.
| Technique | Reduces | Best for | Setup effort |
|---|---|---|---|
| Prompt caching | Cost + latency on repeated prefix | Stable system prompts, long reference docs | Low — mostly prompt ordering |
| Smaller model routing | Cost per token overall | Simple, high-volume tasks | Medium — needs routing logic |
| Output length limits | Cost on generation | Tasks with verbose default output | Low |
| Fine-tuning | Prompt length + occasionally cost per call | Narrow, repeated task types | High |
Notice these aren't mutually exclusive — a well-optimized production pipeline usually stacks a few of them. If you're weighing fine-tuning against just writing a better prompt, our piece on fine-tuning vs. prompting covers that decision directly, and caching is complementary to either choice rather than a replacement for it.
Caching also interacts with how many tokens you're sending in the first place. If your workload depends on stuffing long documents into every call, it's worth reading our comparison of RAG vs. long context — retrieval can shrink what needs to be cached in the first place, and the two techniques compound well together.
For teams tracking spend across providers more broadly, our LLM API pricing breakdown lays out how cached-token discounts factor into each provider's published rate card, which is worth checking before assuming caching savings translate the same way everywhere.
What This Looks Like on a Real Bill
To make the savings concrete, consider a customer-support bot handling 50,000 calls a day, each with a 4,000-token system prompt and tool schema plus a short user message. Uncached, that system-prompt portion alone accounts for the bulk of daily input-token spend, repeated in full on every single call.
With caching enabled and traffic dense enough to keep hitting the cache window, that same 4,000-token block is billed at the discounted cached rate for nearly every call after the first. In the audits we've run for teams at this kind of volume, the input-token line item on the bill dropped by more than half once caching was verified as actually hitting — not just enabled in the SDK, but confirmed via the provider's returned cache-token counts.
The gap between "enabled" and "verified" is exactly where most teams leave savings on the table. A caching flag flipped on in code with no dashboard check behind it is a hope, not a cost optimization.
The Bottom Line
Prompt caching is close to a free win when your workload has a genuinely stable prefix and enough call volume to land inside the cache window — the setup cost is mostly discipline about prompt ordering, not new infrastructure. It's not a fix for a prompt that's simply too long, and it won't help a low-volume, high-variance workload much at all. Check your provider's reported cache-hit token counts before assuming it's working; a silent cache miss looks identical to a normal response.
The Model Drop tracks real inference cost and latency across providers, testing optimization claims like caching against our own benchmark harness rather than provider marketing numbers.