AIInterviewTraining logoAIInterview/Training
ML Infrastructure & GPUs / 06

Explain speculative decoding and the other main levers for cutting LLM generation latency.

Decode runs sequentially and is memory-bound, so latency tricks count. What they grade is whether you can explain speculative decoding's draft-and-verify mechanism (and why it stays exact) along with the other levers and when each fits. This is the latency toolkit.

Updated Sep 2026 · Grounded in real GenAI, LLM, and AI/ML engineering interview loops and written to a senior-engineer editorial bar.

TL;DR: Decoding is sequential and memory-bandwidth-bound: one token per forward pass, each reading the entire model. Speculative decoding has a small fast draft model propose several tokens that the large model verifies in a single parallel pass, keeping the longest correct prefix. The output matches normal decoding exactly, just quicker when acceptance is high. Additional levers: continuous batching and PagedAttention (throughput), quantization (bandwidth), MQA/GQA (cache), shorter prompts and caching.

SPECULATIVE DECODING (press run)
target passes 0tokens 0tokens / pass 0.0
press run to generate
A cheap draft model proposes 4 tokens; the target verifies them all in a single forward pass and accepts the correct prefix, then fixes the first miss. Output quality is untouched, because the target still decides every token. Tokens per target pass is the ceiling on the speedup: the wall-clock gain is lower, since you also pay for the 4 draft steps.

How to approach it. Set out why decode is slow (sequential, memory-bound: you stream the full model per token no matter the batch), then present speculative decoding as a way to verify several tokens per costly pass while staying exact. Finish with the complementary levers so you present the full picture, not a single trick.

A strong answer. Why decode is the bottleneck. Generation is autoregressive: each token needs a full forward pass over the model, and that pass is memory-bandwidth-bound (you read all the weights and the KV cache to produce one token). Latency scales with tokens generated, and you cannot trivially parallelize across the sequence dimension the way prefill does.

Speculative decoding. Pair a small, cheap draft model with the large target model. The draft quickly generates a short run of candidate tokens (say 4). The target then does one forward pass that scores all those positions in parallel and verifies them: it accepts the longest prefix of draft tokens consistent with its own distribution (via a rejection-sampling check) and corrects the first mismatch.

rendering diagram…

If the draft proposed 4 and 3 are accepted, you produced ~3-4 tokens for the cost of one target pass plus the cheap draft passes. Importantly, the output distribution is identical to standard decoding from the target: it is exact, not an approximation, because the target verifies every token. The speedup depends on the acceptance rate (how often the small model agrees with the big one), typically a meaningful multiple on predictable text, less on hard tokens.

Complementary latency and throughput levers:

  • Continuous (in-flight) batching: finished sequences leave and new ones join each step, keeping the GPU busy. The biggest throughput win at serving scale.
  • PagedAttention (vLLM): paged KV-cache storage packs more concurrent requests, indirectly cutting queue latency.
  • Quantization (INT8/FP8/INT4): fewer bytes per weight means less memory traffic per token, directly speeding the bandwidth-bound decode.
  • MQA/GQA: smaller KV cache, more room for batch.
  • Prompt/KV caching and shorter contexts: reuse cached prefixes; trim unnecessary context.
  • Smaller/distilled model or early exit: when the quality budget allows.

The defensible framing: profile first, then combine speculative decoding (latency per request) with continuous batching and quantization (throughput and bandwidth). They address different parts of the cost.

Key takeaways

  • Decode is sequential and bandwidth-bound, so the win is verifying many tokens per expensive target pass, not adding FLOPs.
  • Speculative decoding is exact, not lossy: the target verifies every token, so the output distribution is unchanged.
  • Speedup tracks the draft's acceptance rate; a poor draft can be net-negative because its passes are wasted.
  • Latency levers (speculative decoding) and throughput levers (continuous batching, paging, quantization) stack and target different costs.

What interviewers probe next.

  • "Is speculative decoding lossy?" No. The target verifies, so the output matches what it would have generated alone; you only save time.
  • "What determines the speedup?" The draft's acceptance rate and its relative cost; a good small draft on predictable text gives large gains, a poor one can hurt.
  • "Where does the draft model come from?" A smaller model from the same family, a distilled model, or self-drafting heads (Medusa, EAGLE) that predict several tokens.
  • "Latency vs throughput with batching?" Bigger batches raise throughput but can raise per-request latency; tune to the SLO, and continuous batching softens the tradeoff.

Common mistakes.

  • Calling speculative decoding an approximation; it is exact because the target verifies.
  • Thinking bigger batches always reduce latency; they raise throughput and can raise tail latency.
  • Treating decode as compute-bound and optimizing FLOPs instead of memory bandwidth.
  • Offering one lever; real systems stack speculative decoding, batching, paging, and quantization.
That answer was free, and so are 10 per topic without an account. A free account doubles that to 20, remembers what you have answered, and tracks which topics you are weakest in.no card · Google sign-in · nothing to cancel
HOW DID IT GO?
0
UP NEXT ON YOUR JOURNEY
DISCUSSION · 0

No comments yet — be the first to share your approach.