Platforms

Prompt Caching Explained: Cutting Inference Costs Without Cutting Quality

How prompt caching actually works at the API level, when it saves real money, and the setup mistakes that silently disable it.

Dana Kwon

Contributing Reviewer

Published 6 min read
Woman working on a laptop inside a car, testing sound engineering in an isolated chamber.
Jump to 6 sections

Quick answer: Prompt caching stores the processed representation of a repeated prompt prefix — like a long system prompt or document — so a model provider doesn't re-run the expensive part of inference on every call. It typically cuts cost on the cached portion by 50-90% and cuts latency too, but it only helps when your prompts actually share a stable prefix.

This review covers how prompt caching actually works at the API level, when it saves real money versus when it's a rounding error, and the mistakes that quietly disable it without you noticing. It's for anyone shipping an LLM feature with a system prompt, tool schema, or reference document that repeats across calls.

How Prompt Caching Actually Works

Every call to a large language model starts by processing the input tokens before generating a single output token — this is the "prefill" step, and it scales roughly with prompt length. Prompt caching lets a provider skip re-computing that prefill for tokens it has already processed recently, as long as those tokens appear in the exact same order at the start of the prompt.

At The Model Drop, we've benchmarked cached versus uncached calls across several major providers, and the pattern is consistent: a cache hit on a long system prompt or reference document can cut time-to-first-token by half or more, on top of the direct cost savings on cached input tokens.

The catch is the word "prefix." Caching keys off a shared starting sequence of tokens, not the prompt as a whole. If anything before your reusable content changes between calls — a timestamp, a session ID, a reordered instruction — the cache misses entirely for that call, and you pay full price with no warning in most APIs.

Overhead view of a person analyzing financial documents using a calculator for investment planning.

When Caching Actually Saves Money

Caching pays off when three conditions line up: a long, stable prefix; enough call volume to reuse it within the cache's lifetime; and variable content placed after the cached portion, not mixed into it.

A worked example makes this concrete. Say a support agent sends a 3,000-token system prompt plus tool schema on every call, with only the user's message changing. Without caching, that 3,000-token prefix is billed at full input-token price every single call. With caching, only the first call in a given cache window pays full price; every call after that inside the cache lifetime pays the discounted cached rate on those same 3,000 tokens.

Where caching stops mattering is low-volume, high-variance workloads — a research assistant where every prompt is genuinely different, or an endpoint that gets one call every few minutes, well outside typical cache lifetimes (usually 5-10 minutes across providers, though this varies and changes over time). In our own cost audits, we've seen teams add caching to a low-traffic internal tool and see almost no bill impact, because the cache had already expired by the time the next call came in.

The underlying mechanics trace back to how transformer inference splits into a prefill phase and a decode phase — a distinction covered in the original attention-mechanism research indexed on arXiv. Prefill is the parallelizable, cache-friendly part; decode, which generates tokens one at a time, is not something caching speeds up.

The Setup Mistakes That Silently Disable It

Caching failures are usually invisible unless you're checking cache-hit metrics directly, since an uncached call still returns a correct answer — it's just full price.

  1. Variable content before the cacheable prefix. A timestamp, request ID, or user name placed at the very top of the prompt breaks the match for everything after it. Move stable content first.
  2. Non-deterministic prompt construction. If your prompt-building code reorders a list of tools or documents differently on each call (e.g. from an unsorted dictionary), the token sequence changes even though the content is logically the same.
  3. Cache lifetime mismatches. Sending cacheable calls further apart than the provider's cache TTL means every call is effectively a fresh cache write, not a hit.
  4. Not reading provider-reported cache token counts. Most APIs return how many tokens were served from cache per response — ignoring that field means shipping a caching setup you can't verify is actually working.

IEEE Spectrum has reported on how quickly inference costs can scale for teams that don't audit these details, since a broken cache doesn't throw an error — it just quietly bills every call at the uncached rate until someone checks the dashboard.

Focused view of a blue illuminated laptop keyboard emphasizing keys and technology.

Caching vs. Other Cost Levers

Prompt caching is one of several cost levers, and it's worth knowing where it fits relative to the others.

Caching vs. Other Cost Levers
TechniqueReducesBest forSetup effort
Prompt cachingCost + latency on repeated prefixStable system prompts, long reference docsLow — mostly prompt ordering
Smaller model routingCost per token overallSimple, high-volume tasksMedium — needs routing logic
Output length limitsCost on generationTasks with verbose default outputLow
Fine-tuningPrompt length + occasionally cost per callNarrow, repeated task typesHigh

Notice these aren't mutually exclusive — a well-optimized production pipeline usually stacks a few of them. If you're weighing fine-tuning against just writing a better prompt, our piece on fine-tuning vs. prompting covers that decision directly, and caching is complementary to either choice rather than a replacement for it.

Caching also interacts with how many tokens you're sending in the first place. If your workload depends on stuffing long documents into every call, it's worth reading our comparison of RAG vs. long context — retrieval can shrink what needs to be cached in the first place, and the two techniques compound well together.

For teams tracking spend across providers more broadly, our LLM API pricing breakdown lays out how cached-token discounts factor into each provider's published rate card, which is worth checking before assuming caching savings translate the same way everywhere.

What This Looks Like on a Real Bill

To make the savings concrete, consider a customer-support bot handling 50,000 calls a day, each with a 4,000-token system prompt and tool schema plus a short user message. Uncached, that system-prompt portion alone accounts for the bulk of daily input-token spend, repeated in full on every single call.

With caching enabled and traffic dense enough to keep hitting the cache window, that same 4,000-token block is billed at the discounted cached rate for nearly every call after the first. In the audits we've run for teams at this kind of volume, the input-token line item on the bill dropped by more than half once caching was verified as actually hitting — not just enabled in the SDK, but confirmed via the provider's returned cache-token counts.

The gap between "enabled" and "verified" is exactly where most teams leave savings on the table. A caching flag flipped on in code with no dashboard check behind it is a hope, not a cost optimization.

The Bottom Line

Prompt caching is close to a free win when your workload has a genuinely stable prefix and enough call volume to land inside the cache window — the setup cost is mostly discipline about prompt ordering, not new infrastructure. It's not a fix for a prompt that's simply too long, and it won't help a low-volume, high-variance workload much at all. Check your provider's reported cache-hit token counts before assuming it's working; a silent cache miss looks identical to a normal response.

The Model Drop tracks real inference cost and latency across providers, testing optimization claims like caching against our own benchmark harness rather than provider marketing numbers.

How much does prompt caching actually save on API costs?
Cached input tokens typically cost 50-90% less than standard input tokens, depending on the provider. The savings apply only to the cached prefix portion of a prompt, so total bill impact depends on how much of your prompt is stable versus variable.
How long does a prompt cache last before it expires?
Most providers keep a cache alive for roughly 5-10 minutes of inactivity, though exact windows vary by provider and change over time. Calls spaced further apart than the cache lifetime will not hit the cache and are billed at the standard rate.
Does prompt caching reduce response latency, not just cost?
Yes. Skipping the prefill step on cached tokens reduces time-to-first-token noticeably, often cutting it by half or more on long cached prefixes, in addition to the direct per-token cost savings.
Why would prompt caching stop working without any error message?
Caching fails silently when the token sequence before the cached content changes between calls — a timestamp, reordered list, or variable inserted too early all break the prefix match. The call still succeeds; it is just billed at full price.

Written by

Dana Kwon

Contributing Reviewer

Dana builds and re-runs a fixed evaluation harness against every model and coding tool Model Drop reviews, so a rating from March means the same thing as a rating in August.

Covers

  • model evaluation
  • inference infrastructure
  • cost & performance comparisons