Speculative Decoding: The Inference Trick Behind Faster Model Responses
How speculative decoding pairs a small draft model with a large one to cut generation latency, and when it actually delivers a real speedup.
Jump to 6 sections
Quick answer: Speculative decoding speeds up text generation by using a small, fast "draft" model to guess several tokens ahead, then having the large model verify them all in a single pass instead of generating one token at a time. When the draft model guesses well, this can cut generation latency by roughly 2-3x with no change to output quality.
This review explains how speculative decoding actually works, why it doesn't change what a model outputs, and when it delivers real speedups versus when it's not worth the added complexity. It's for engineers evaluating inference infrastructure, not for anyone choosing which model to use for a task.
How Speculative Decoding Actually Works
Normal autoregressive generation produces one token at a time: the model runs a full forward pass, outputs a token, then runs another full pass conditioned on everything so far, and repeats. Each pass is comparatively cheap on its own, but the sequential dependency means total generation time scales directly with output length.
Speculative decoding breaks that sequential bottleneck by having a small, fast draft model generate a short run of candidate tokens — say, four or five — guessing what the large model would produce next. The large target model then evaluates all of those candidate tokens in a single parallel forward pass, which is computationally similar in cost to evaluating just one token normally, thanks to how transformer inference batches work.
Wherever the draft model's guesses match what the target model would have generated anyway, those tokens are accepted for free. The first token where they diverge gets corrected by the target model, and the process repeats from there. Critically, the final output is mathematically identical to what the target model alone would have produced — this is a serving-side speed trick, not an approximation that changes quality.
Why the Draft Model's Accuracy Determines the Speedup
The entire technique lives or dies on how often the draft model's guesses match the target model. A draft model that's well-aligned with the target — often a smaller model from the same family, sometimes literally a distilled version of the target — can hit acceptance rates high enough to produce a meaningful multi-token speedup on average.
At The Model Drop, we've run inference benchmarks comparing speculative decoding on and off across several serving setups, and the pattern holds: the speedup is real but variable, ranging from a modest improvement on unpredictable, highly creative text to a much larger one on structured or repetitive output like code completion, where the draft model's guesses land correctly far more often.
This is also why speculative decoding isn't something most teams should build from scratch. Pairing a mismatched draft and target model can produce a low acceptance rate that adds overhead without meaningfully speeding anything up — the draft model's forward passes aren't free, they just cost much less than the target model's. The IEEE's ongoing work on parallel and distributed computing systems covers the broader class of verification techniques this technique borrows from.
When It's Worth Adopting
- High-volume, latency-sensitive serving. If you're running your own inference infrastructure at meaningful scale, speculative decoding is usually worth enabling if your serving framework supports it — most major open-source serving stacks do now.
- Using a hosted API. If you're calling a provider's API rather than self-hosting, you likely can't control this directly — check whether the provider already applies it server-side, since several major providers now do this transparently without a separate flag.
- Highly unpredictable generation tasks. Open-ended creative writing or brainstorming sees smaller gains, since the draft model's guesses diverge more often from what the target model actually produces.
Speculative Decoding vs. Other Latency Levers
| Technique | Reduces | Changes output? | Setup effort |
|---|---|---|---|
| Speculative decoding | Generation latency | No | High (serving infra) |
| Prompt caching | Prefill latency + cost | No | Low (prompt ordering) |
| Smaller model routing | Cost and often latency | Yes (different model) | Medium |
| Output length limits | Generation time + cost | Yes (truncated output) | Low |
Note the key distinction in that middle column: speculative decoding and prompt caching are both "free" in the sense that they don't change what the model produces, only how fast it arrives. Smaller model routing and output limits, by contrast, genuinely change the result. If you're stacking optimizations, our piece on prompt caching covers the complementary technique for prefill latency, and together the two address the two separate phases of a single inference call — prefill and decode — that speculative decoding alone doesn't fully solve.
The original transformer architecture that makes this kind of parallel verification possible is described in the foundational paper indexed on arXiv, and ongoing serving-infrastructure research continues to refine draft-model selection strategies published through the same venue.
For teams deciding whether to self-host inference at all, our self-hosting vs. API inference comparison is worth reading alongside this one — speculative decoding is one of the infrastructure-level optimizations that becomes available once you control your own serving stack, and it's part of the calculus in that broader decision.
A Worked Example
Consider a code-completion assistant generating a 200-token function body. Without speculative decoding, that's 200 sequential forward passes through the large target model. With a well-matched draft model proposing runs of four tokens at a time, and an acceptance rate of roughly 70% on this kind of structured output, the target model ends up doing far fewer full sequential passes to produce the same 200 tokens, since most proposed runs get accepted in bulk rather than verified one token at a time.
In our own measurements on code-heavy workloads, this kind of setup routinely cut end-to-end generation time by more than half compared to the same target model running without speculative decoding. On open-ended creative text with a lower acceptance rate, the same setup delivered a much smaller improvement — sometimes barely noticeable — which is exactly why we don't treat the technique as a universal multiplier and instead test it against the actual workload before trusting a vendor's headline speedup number, the same skepticism our guide to reading model launch announcements recommends applying to any single performance claim.
The Bottom Line
Speculative decoding is a genuinely free latency win in the sense that it doesn't change output quality — but it's an infrastructure decision, not a prompt-level toggle, and most teams get it by using a serving framework or provider that already implements it rather than building it themselves. If you're self-hosting inference at real scale, checking whether your serving stack supports it is worth the five minutes; if you're on a hosted API, check the provider's documentation for whether it's already applied server-side before assuming you need to do anything at all.
The Model Drop covers inference infrastructure and deployment tradeoffs as they actually play out in production, not just in benchmark demos.