Speculative Decoding Explained: Speed Up Local LLM Inference

September 29, 202613 min readSam Austin
Contents

Speculative decoding speeding up local LLM inference on a GPU
Speculative decoding speeding up local LLM inference on a GPU

Figure 1: The compute you're already paying for but not using

Your GPU sits mostly idle during LLM generation, waiting on memory bandwidth rather than compute. Speculative decoding exploits exactly that wasted capacity, letting a small fast model guess ahead while the big model verifies — often for a genuinely free 2-3x speedup with zero quality loss. Let's break down how it actually works and get it running.

I was skeptical the first time I heard "faster generation, same output quality, no catch." Turns out the trick is legitimately clever, not a marketing claim.

The Bottleneck This Solves

Autoregressive generation, the normal way LLMs produce text, generates one token at a time, and each token requires a full forward pass through the model. Here's the counterintuitive part: that forward pass is memory-bandwidth-bound, not compute-bound. Your GPU's compute units sit mostly idle while data streams from memory, because generating one token doesn't use anywhere near your hardware's full compute capacity.

Speculative decoding exploits that unused capacity directly. Instead of computing one token per forward pass, you can verify several proposed tokens in a single forward pass for roughly the same cost, since the bottleneck was never compute in the first place.

If you've never profiled where your generation time actually goes, our local LLM benchmarking guide walks through measuring tokens per second before you start optimizing.

The Core Mechanism

Speculative decoding is a collaboration between two models:

  • The target model: Your actual model, the one whose output quality you want (e.g., Llama-3.3-70B)
  • The draft model: A smaller, faster model that proposes candidate tokens (e.g., a 1B model, or something even lighter)

The cycle works like this:

  1. Draft phase: The small draft model generates K speculative tokens quickly, autoregressively
  2. Verification phase: The target model scores the entire draft sequence in one batched forward pass
  3. Accept or reject: Tokens matching what the target model would have generated get accepted; the first mismatch gets corrected, and everything after it is discarded
  4. Repeat from the first rejected token

The critical guarantee here: verification produces mathematically identical output to standard decoding. You're not trading quality for speed. The target model still makes every final decision; the draft model just proposes candidates that get checked, not trusted blindly.

Ever wondered why this doesn't just introduce a smaller model's mistakes into your output? That's the answer: rejection during verification catches every deviation.

How Much Speedup Should You Expect?

Real numbers vary a lot by method and workload, but here's a grounded picture from recent production benchmarks:

MethodTypical SpeedupDraft Model Needed
Draft model (classic)1.5–2xYes, separate small model
N-gram / prompt lookupModest, workload-dependentNo, pattern matching only
Medusa2.2–3.6xNo, extra heads on target model
EAGLE / EAGLE-22.5–2.8x typicalYes, lightweight 1-layer draft
EAGLE-33–6.5xYes, pre-trained heads
P-EAGLEUp to 4–5x on coding tasksYes, parallel drafting variant

Coding-heavy workloads see the largest gains, with the EAGLE-3 paper reporting close to 4.8x on HumanEval for a 70B Llama model. Speedups shrink meaningfully as batch size grows. At batch size 1 (a single user chatting), EAGLE-style methods hit their highest multipliers; in heavily batched serving scenarios, the GPU is already busy with other requests, so there's less idle compute left to exploit.

The Methods, Explained Simply

Draft Model Speculation (Classic)

The original approach: pair your target model with a genuinely separate, smaller model of similar architecture. Both run autoregressively, but the draft model is cheap enough that running it several times costs less than one target model forward pass.

vllm serve meta-llama/Llama-3.1-70B-Instruct \
  --speculative-model meta-llama/Llama-3.1-8B-Instruct \
  --num-speculative-tokens 4

Simple to set up, since you're just pointing at two existing models. The trade-off is memory: you need VRAM for both models simultaneously, and a draft model in FP16 or FP8 typically adds 2 to 6 GB depending on its size.

N-gram / Prompt Lookup Decoding

No neural draft model at all. This method identifies repeated patterns in the prompt or generation history and proposes tokens based on observed sequences:

vllm serve meta-llama/Llama-3.1-70B-Instruct \
  --speculative-model "[ngram]" \
  --ngram-prompt-lookup-max 4 \
  --num-speculative-tokens 4

This works surprisingly well for tasks with lots of repetition — code editing, document summarization, anything where the output echoes parts of the input. It needs zero extra training or extra model weights, making it the easiest method to try first.

Medusa

Medusa takes a fundamentally different approach by not using a separate draft model at all. It attaches multiple extra prediction heads directly onto the target model, letting it forecast several tokens ahead in a single forward pass. With K Medusa heads, the model predicts up to K tokens simultaneously.

The trade-off: Medusa requires model modification and retraining to add those heads, so it's not a drop-in addition to an existing model the way a draft model or n-gram approach is. Also worth knowing: compatibility issues are real. Some newer model architectures lack the specific interface Medusa's implementation requires, so verify support before committing to this path.

EAGLE and Its Variants

EAGLE sits architecturally between Medusa and classic draft models. A tiny single-layer transformer runs autoregressively, but instead of working from raw tokens like a classic draft model, it works from the target model's own hidden states. This gives EAGLE higher acceptance rates than Medusa and lower draft cost than a full separate draft model, since the draft component is just one layer deep.

EAGLE has become the go-to for serious deployments. It's currently the top performer on the Spec-Bench leaderboard, with roughly 80% draft acceptance and 2.5–2.8x typical speedups. EAGLE-3 pushed further with multi-layer feature aggregation, reaching 3–6.5x speedups in some benchmarks. P-EAGLE, contributed to mainline vLLM by AWS in early 2026, extends this further with parallel drafting: instead of generating draft tokens one at a time, the draft model produces all K candidates in a single forward pass, then builds a verification tree from them, delivering up to 1.69x speedup over vanilla EAGLE-3 on real workloads.

Setting Up Speculative Decoding in Practice

With vLLM

vllm serve meta-llama/Llama-3.1-8B-Instruct \
  --dtype bfloat16 \
  --speculative-config '{"model": "yuhuili/EAGLE-LLaMA3-Instruct-8B", "num_speculative_tokens": 5}'

EAGLE and EAGLE-3 need specific pre-trained checkpoints matching your target model, unlike n-gram methods which need nothing extra. Check HuggingFace's Speculative Decoding Modules collection or your target model's community for a matching draft checkpoint before assuming one exists.

With llama.cpp

For local, consumer-hardware use, llama.cpp supports classic draft-model speculation directly:

llama-server -m target-model.gguf \
  --model-draft draft-model.gguf \
  --draft-max 16

Pick a draft model that's architecturally similar to your target (same tokenizer family matters a lot here) and meaningfully smaller, so it's cheap to run multiple speculative steps. Our llama.cpp tutorial covers the basics first if you're still setting up.

A Practical Decision Framework

Let me save you some research time with straightforward guidance:

  1. Repetitive workload, code editing, summarization? N-gram / prompt lookup first — zero extra weights, minutes to test
  2. Single user on local hardware, 70B-class target? Classic draft model via --model-draft, quantize it aggressively
  3. Already on vLLM with a matching checkpoint? EAGLE — it's the current leaderboard leader for a reason
  4. Serving many concurrent users? Benchmark before assuming; batch size and hardware choice decide how much idle compute you actually have
  5. Small target model (7B or less)? Expect thin gains — the draft overhead eats the win
  6. Need smaller models in general? Knowledge distillation and quantization attack the same problem from the other side

The mistake isn't picking the wrong method — it's deploying stacked speculative tricks that quietly fall back to one working method under the hood.

Common Mistakes People Make

Mismatched tokenizers

Recall the llama.cpp setup — a mismatched tokenizer between draft and target model breaks the whole scheme, so stay in the same model family.

Oversized draft model

Recall the classic method — keep the draft small relative to the target, or the extra forward passes cancel out the batching win.

Ignoring batch size

Recall the benchmark section — speedups that look great at batch size 1 often shrink badly under real serving load.

Expecting free wins on tiny targets

Recall the tuning notes — draft-model methods work best when the target is large; on a 7B the overhead can eat the gains.

Assuming stacked methods multiply

FYI: combining multiple speculative techniques doesn't reliably compose — one detailed public writeup found that stacking four different approaches fell back to just one working method due to compatibility issues. Test your specific combination rather than assuming.

  • Computer Architecture: A Quantitative Approach by John L. Hennessy and David A. Patterson — the canonical explanation of why memory bandwidth, not FLOPs, is the wall you're hitting here.
  • Programming Massively Parallel Processors by David B. Kirk and Wen-mei W. Hwu — how GPU kernels actually behave, which makes the "compute sits idle" claim obvious rather than surprising.
  • Systems Performance by Brendan Gregg — the discipline of measuring before optimizing, applied here to your own inference stack instead of a datacenter.

Want to Go Deeper?

If you want structured practice on ML systems and serving, Educative's ML courses include hands-on labs that pair well with this kind of performance work. The unlimited plan is useful when you're working through several optimization techniques in one stretch.

Unlock AI That Actually Works

Get lifetime access to GPT-6 Astra, Claude Fable 5.1, Gemini 3.5, Grok 4.5, and more — all in one platform. Build websites, apps, videos, content, and digital products from a single command. No monthly fees. No tool-hopping.

Click here to get GPTAstra Max now — one-time payment, lifetime access.

Frequently Asked Questions

What is speculative decoding?

A technique where a small fast draft model proposes several candidate tokens, then the large target model verifies them all in one batched forward pass. Accepted tokens are kept, the first mismatch is corrected, and everything after it is discarded. The target model still makes every final decision.

Does speculative decoding change output quality?

No. Verification produces mathematically identical output to standard decoding, because every proposed token is checked against what the target model itself would have produced. You get speed without trading away quality or introducing the draft model's mistakes.

How much speedup can I expect?

Classic draft-model speculation typically lands at 1.5-2x, Medusa at 2.2-3.6x, and EAGLE-family methods at 2.5-2.8x typical with peaks of 3-6.5x. Coding-heavy workloads gain the most, and gains shrink as batch size grows because the GPU is already busy.

Do I need a separate draft model?

Not always. N-gram or prompt-lookup decoding needs no extra model at all, and Medusa adds prediction heads directly to the target model. Classic draft models and EAGLE variants do need a second component, with EAGLE requiring a checkpoint that matches your specific target model.

Which speculative decoding method should I use first?

Start with n-gram or prompt-lookup decoding if your workload is repetitive, since it costs nothing to set up. Move to a draft model or EAGLE when you need real speedups on open-ended generation and are willing to manage an extra checkpoint.

Does batch size affect speculative decoding?

Yes, meaningfully. At batch size 1, a single user chatting, idle compute is abundant and speedups are highest. In heavily batched serving the GPU is already occupied, so there is less free capacity to exploit and multipliers drop.

How do I enable it in llama.cpp or vLLM?

In llama.cpp, pass a draft GGUF with --model-draft and control depth with --draft-max. In vLLM, either set --speculative-model for classic and n-gram approaches, or use --speculative-config with an EAGLE checkpoint plus num_speculative_tokens.

Wrapping This Up

Speculative decoding delivers a genuinely rare thing in ML systems: meaningfully faster generation with mathematically identical output quality, by exploiting idle compute capacity that standard autoregressive decoding leaves on the table. Draft models, n-gram matching, Medusa heads, and EAGLE variants all attack the same bottleneck from different angles, with EAGLE-family methods currently leading on raw speedup.

Will it help every workload equally? No, batch-1 chat sees the biggest wins, and heavily batched serving sees smaller gains as GPU idle time shrinks. But for local single-user inference specifically, where you're often waiting on one response at a time, this is close to a free lunch. Pair a small draft model with your local llama.cpp setup tonight and watch your tokens-per-second jump without touching a single quality setting.

Once you've measured the gain, compare it against what model quantization and GGUF file selection buy you — the three stack nicely on local hardware.