Platforms

Speculative Decoding: The Inference Trick Behind Faster Model Responses

How speculative decoding pairs a small draft model with a large one to cut generation latency, and when it actually delivers a real speedup.

Dana Kwon

Contributing Reviewer

Published 6 min read
Mechanic focusing on car engine repair, showcasing skill and precision.
Jump to 6 sections

Quick answer: Speculative decoding speeds up text generation by using a small, fast "draft" model to guess several tokens ahead, then having the large model verify them all in a single pass instead of generating one token at a time. When the draft model guesses well, this can cut generation latency by roughly 2-3x with no change to output quality.

This review explains how speculative decoding actually works, why it doesn't change what a model outputs, and when it delivers real speedups versus when it's not worth the added complexity. It's for engineers evaluating inference infrastructure, not for anyone choosing which model to use for a task.

How Speculative Decoding Actually Works

Normal autoregressive generation produces one token at a time: the model runs a full forward pass, outputs a token, then runs another full pass conditioned on everything so far, and repeats. Each pass is comparatively cheap on its own, but the sequential dependency means total generation time scales directly with output length.

Speculative decoding breaks that sequential bottleneck by having a small, fast draft model generate a short run of candidate tokens — say, four or five — guessing what the large model would produce next. The large target model then evaluates all of those candidate tokens in a single parallel forward pass, which is computationally similar in cost to evaluating just one token normally, thanks to how transformer inference batches work.

Wherever the draft model's guesses match what the target model would have generated anyway, those tokens are accepted for free. The first token where they diverge gets corrected by the target model, and the process repeats from there. Critically, the final output is mathematically identical to what the target model alone would have produced — this is a serving-side speed trick, not an approximation that changes quality.

Contemporary computer with black screen placed on stand near row of server steel racks in data center

Why the Draft Model's Accuracy Determines the Speedup

The entire technique lives or dies on how often the draft model's guesses match the target model. A draft model that's well-aligned with the target — often a smaller model from the same family, sometimes literally a distilled version of the target — can hit acceptance rates high enough to produce a meaningful multi-token speedup on average.

At The Model Drop, we've run inference benchmarks comparing speculative decoding on and off across several serving setups, and the pattern holds: the speedup is real but variable, ranging from a modest improvement on unpredictable, highly creative text to a much larger one on structured or repetitive output like code completion, where the draft model's guesses land correctly far more often.

This is also why speculative decoding isn't something most teams should build from scratch. Pairing a mismatched draft and target model can produce a low acceptance rate that adds overhead without meaningfully speeding anything up — the draft model's forward passes aren't free, they just cost much less than the target model's. The IEEE's ongoing work on parallel and distributed computing systems covers the broader class of verification techniques this technique borrows from.

When It's Worth Adopting

  1. High-volume, latency-sensitive serving. If you're running your own inference infrastructure at meaningful scale, speculative decoding is usually worth enabling if your serving framework supports it — most major open-source serving stacks do now.
  2. Using a hosted API. If you're calling a provider's API rather than self-hosting, you likely can't control this directly — check whether the provider already applies it server-side, since several major providers now do this transparently without a separate flag.
  3. Highly unpredictable generation tasks. Open-ended creative writing or brainstorming sees smaller gains, since the draft model's guesses diverge more often from what the target model actually produces.
Dynamic blurred image from a vehicle's windshield capturing a rural road in motion.

Speculative Decoding vs. Other Latency Levers

Speculative Decoding vs. Other Latency Levers
TechniqueReducesChanges output?Setup effort
Speculative decodingGeneration latencyNoHigh (serving infra)
Prompt cachingPrefill latency + costNoLow (prompt ordering)
Smaller model routingCost and often latencyYes (different model)Medium
Output length limitsGeneration time + costYes (truncated output)Low

Note the key distinction in that middle column: speculative decoding and prompt caching are both "free" in the sense that they don't change what the model produces, only how fast it arrives. Smaller model routing and output limits, by contrast, genuinely change the result. If you're stacking optimizations, our piece on prompt caching covers the complementary technique for prefill latency, and together the two address the two separate phases of a single inference call — prefill and decode — that speculative decoding alone doesn't fully solve.

The original transformer architecture that makes this kind of parallel verification possible is described in the foundational paper indexed on arXiv, and ongoing serving-infrastructure research continues to refine draft-model selection strategies published through the same venue.

For teams deciding whether to self-host inference at all, our self-hosting vs. API inference comparison is worth reading alongside this one — speculative decoding is one of the infrastructure-level optimizations that becomes available once you control your own serving stack, and it's part of the calculus in that broader decision.

A Worked Example

Consider a code-completion assistant generating a 200-token function body. Without speculative decoding, that's 200 sequential forward passes through the large target model. With a well-matched draft model proposing runs of four tokens at a time, and an acceptance rate of roughly 70% on this kind of structured output, the target model ends up doing far fewer full sequential passes to produce the same 200 tokens, since most proposed runs get accepted in bulk rather than verified one token at a time.

In our own measurements on code-heavy workloads, this kind of setup routinely cut end-to-end generation time by more than half compared to the same target model running without speculative decoding. On open-ended creative text with a lower acceptance rate, the same setup delivered a much smaller improvement — sometimes barely noticeable — which is exactly why we don't treat the technique as a universal multiplier and instead test it against the actual workload before trusting a vendor's headline speedup number, the same skepticism our guide to reading model launch announcements recommends applying to any single performance claim.

The Bottom Line

Speculative decoding is a genuinely free latency win in the sense that it doesn't change output quality — but it's an infrastructure decision, not a prompt-level toggle, and most teams get it by using a serving framework or provider that already implements it rather than building it themselves. If you're self-hosting inference at real scale, checking whether your serving stack supports it is worth the five minutes; if you're on a hosted API, check the provider's documentation for whether it's already applied server-side before assuming you need to do anything at all.

The Model Drop covers inference infrastructure and deployment tradeoffs as they actually play out in production, not just in benchmark demos.

Does speculative decoding change the quality of a model’s output?
No. The final output is mathematically identical to what the target model would produce on its own. Speculative decoding only changes how fast the tokens arrive, using a small draft model to propose candidates the large model verifies in parallel.
How much faster is generation with speculative decoding enabled?
Speedups commonly land around 2-3x, but this varies significantly with how well the draft model’s guesses match the target model. Predictable text like code sees larger gains than open-ended creative writing.
Can I enable speculative decoding when using a hosted AI API?
Usually not directly, since it is a serving-infrastructure optimization. Several major hosted providers apply it transparently on their end already; check the provider’s documentation rather than assuming you need to configure anything.
Is speculative decoding worth building for a self-hosted inference setup?
For high-volume, latency-sensitive serving, yes, and most major open-source serving frameworks now support it out of the box. Building a custom implementation from scratch is rarely worth it compared to using an existing serving stack’s support.

Written by

Dana Kwon

Contributing Reviewer

Dana builds and re-runs a fixed evaluation harness against every model and coding tool Model Drop reviews, so a rating from March means the same thing as a rating in August.

Covers

  • model evaluation
  • inference infrastructure
  • cost & performance comparisons