syedazeez337/speculative-decoding-explained

A beginner-friendly, from-scratch course on speculative decoding: plain-English lessons, an interactive explainer, designed PDF notes, and runnable cross-platform code (CUDA / Apple Metal / CPU).

★ 0Forks 0TypstGitHub ↗Compare
educationllmllm-inferencemachine-learningpytorchspeculative-decodingtransformerstypst

README

Speculative Decoding, Explained

Learn speculative decoding — the trick that makes large language models generate text 2–3× faster without changing their output — from zero, with plain-English lessons, an interactive page, and code you can actually run.

In one sentence: a small, fast draft model guesses the next few words; the big, accurate target model checks all the guesses in a single pass and fixes any mistakes. You get the big model's exact quality, just faster.


Table of contents


What's inside

Folder / file What it is Need to code?
guide/ A 4-chapter written course for absolute beginners. Start here to learn. No
notes/ The same course as beautifully designed, print-ready PDFs (typeset math, diagrams, colour-coded callouts). No
speculative_decoding_explainer.html An interactive page — watch the loop, drag the bars, move the sliders. Just open it in a browser. No
reference/spec_decode_numpy.py The algorithm in tiny, dependency-free code. Proves the "no quality loss" claim. A little
reference/spec_decode_pytorch.py The same thing with real language models. Runs on any computer. A little
research/ Notes on how the best explainers teach this, plus the design plan. No
papers/ The six source papers (arXiv links). No

Requirements

  • Python 3.9 or newer. Check with python --version. Don't have it? Get it from python.org/downloads.
  • That's all you need for option A. Options B needs two extra packages (the commands below install them).

First, get the code. Clone it (or download the ZIP from your host and unzip), then move into the folder. Every command below is run from inside this folder.

git clone <your-repo-url> speculative-decoding
cd speculative-decoding

Recommended (optional): work inside a virtual environment so nothing installs globally.

python -m venv .venv
# Windows:        .venv\Scripts\activate
# macOS / Linux:  source .venv/bin/activate

Quick start

Pick the path that fits you. They're independent — you can do any one first.

A. See it work (1 minute, no GPU)

The fastest way to get it. This runs the algorithm on tiny made-up numbers and proves the output is identical to the big model's.

pip install -r requirements.txt
python reference/spec_decode_numpy.py

You should see something like this — note that the measured output matches the target exactly, even though the draft disagrees:

--- Losslessness check (200k trials) ---
  token   target p  measured  match
  A          0.300     0.301   ############
  B          0.400     0.400   ################
  C          0.300     0.299   ############
  max error = 0.0010  ->  PASS: output matches target

That PASS is the whole point: faster, but provably no loss in quality.

B. Run it with real AI models

Watch speculative decoding drive actual language models. Works on Windows, macOS, and Linux — it automatically uses your GPU if you have one.

pip install -r requirements-pytorch.txt
python reference/spec_decode_pytorch.py

The first run downloads the models (a few hundred MB), so give it a minute. Prefer smaller/faster models? Use:

python reference/spec_decode_pytorch.py --draft distilgpt2 --target gpt2

You should see the generated text plus a speed stat:

Device: cuda
--- OUTPUT ---
The key idea behind speculative decoding is ...
--- STATS ---
avg tokens per target pass: 3.23   (>1 means speedup)

avg tokens per target pass: 3.23 means the big model ran once for every ~3 words instead of once per word — that's the speedup, live.

C. Open the interactive page

No install, no terminal. Just double-click speculative_decoding_explainer.html (or open it in any browser). It's fully self-contained. Drag the probability bars and move the sliders to build intuition.


Which hardware it uses

The real-model demo (option B) picks the best available accelerator for you — you never change the code:

Your computer What it uses automatically
NVIDIA GPU (Windows / Linux) CUDA
Apple Silicon Mac (M1/M2/M3/M4) Apple Metal (MPS)
No GPU / anything else CPU (slower, but always works)

Want to force one? Add --device cpu (or cuda, or mps). On Apple Silicon, CPU-fallback for any unsupported operation is enabled automatically, so it "just works."


Command options

All options for the real-model demo (spec_decode_pytorch.py):

Option Default Meaning
--draft distilgpt2 The small "guesser" model (Hugging Face name).
--target gpt2-large The big "checker" model.
--prompt "The key idea…" The starting text.
--gamma 4 How many words to guess per round. Try 3–7.
--max-new-tokens 60 How many words to generate.
--temperature 1.0 Randomness. 0 = most likely word every time.
--device auto auto, cuda, mps, or cpu.
--seed 0 Change for different random output.

Example — a longer, more focused generation on the GPU:

python reference/spec_decode_pytorch.py \
  --draft distilgpt2 --target gpt2 \
  --prompt "Once upon a time" --gamma 5 \
  --max-new-tokens 100 --temperature 0.7

Troubleshooting

Problem Fix
python: command not found Try python3 instead of python. On Windows, reinstall Python and tick "Add Python to PATH".
No module named numpy Run pip install -r requirements.txt (and make sure your virtual env is activated).
No module named torch That's only for option B — run pip install -r requirements-pytorch.txt.
Model download is slow / fails It's a one-time download; retry on a stable connection. Use smaller models: --draft distilgpt2 --target gpt2.
Out of memory on GPU Use smaller models (--target gpt2) or force CPU with --device cpu.
avg tokens per target pass is low (~1) Normal if draft and target rarely agree. Use a draft from the same family as the target for higher agreement.
pip installs to the wrong Python Use python -m pip install ... so pip matches the Python you're running.

Run the tests

Confirms the two core claims — losslessness and the speedup math:

# with pytest:
pip install pytest
pytest

# or with no extra install:
python tests/test_losslessness.py

Expected: 6 passed.


Learn the concepts

New to the topic? Read the course in order (about an hour total):

  1. guide/00-intuition.md — the whole idea, no math
  2. guide/01-math.md — the accept/reject rule and why it's exact
  3. guide/02-algorithm.md — the step-by-step recipe + the code
  4. guide/03-methods.md — the modern methods (token trees, Medusa, EAGLE)

Prefer PDFs? The same four chapters are in notes/ as designed, print-ready documents (built from Typst source — rebuild with typst compile).

Notation used everywhere: p = the target (big/accurate) model · q = the draft (small/fast) model · γ (gamma) = how many words the draft guesses before the target checks.


License

MIT © 2026 Syed Azeez. The PDFs in papers/ belong to their respective authors — see papers/README.md for arXiv links.

Contributors

syedazeez337

Issues