Learn speculative decoding — the trick that makes large language models generate text 2–3× faster without changing their output — from zero, with plain-English lessons, an interactive page, and code you can actually run.
In one sentence: a small, fast draft model guesses the next few words; the big, accurate target model checks all the guesses in a single pass and fixes any mistakes. You get the big model's exact quality, just faster.
- What's inside
- Requirements
- Quick start — three ways to begin
- Which hardware it uses
- Command options
- Troubleshooting
- Run the tests
- Learn the concepts
- License
| Folder / file | What it is | Need to code? |
|---|---|---|
guide/ |
A 4-chapter written course for absolute beginners. Start here to learn. | No |
notes/ |
The same course as beautifully designed, print-ready PDFs (typeset math, diagrams, colour-coded callouts). | No |
speculative_decoding_explainer.html |
An interactive page — watch the loop, drag the bars, move the sliders. Just open it in a browser. | No |
reference/spec_decode_numpy.py |
The algorithm in tiny, dependency-free code. Proves the "no quality loss" claim. | A little |
reference/spec_decode_pytorch.py |
The same thing with real language models. Runs on any computer. | A little |
research/ |
Notes on how the best explainers teach this, plus the design plan. | No |
papers/ |
The six source papers (arXiv links). | No |
- Python 3.9 or newer. Check with
python --version. Don't have it? Get it from python.org/downloads. - That's all you need for option A. Options B needs two extra packages (the commands below install them).
First, get the code. Clone it (or download the ZIP from your host and unzip), then move into the folder. Every command below is run from inside this folder.
git clone <your-repo-url> speculative-decoding cd speculative-decoding
Recommended (optional): work inside a virtual environment so nothing installs globally.
python -m venv .venv # Windows: .venv\Scripts\activate # macOS / Linux: source .venv/bin/activate
Pick the path that fits you. They're independent — you can do any one first.
The fastest way to get it. This runs the algorithm on tiny made-up numbers and proves the output is identical to the big model's.
pip install -r requirements.txt
python reference/spec_decode_numpy.pyYou should see something like this — note that the measured output matches the target exactly, even though the draft disagrees:
--- Losslessness check (200k trials) ---
token target p measured match
A 0.300 0.301 ############
B 0.400 0.400 ################
C 0.300 0.299 ############
max error = 0.0010 -> PASS: output matches target
That PASS is the whole point: faster, but provably no loss in quality.
Watch speculative decoding drive actual language models. Works on Windows, macOS, and Linux — it automatically uses your GPU if you have one.
pip install -r requirements-pytorch.txt
python reference/spec_decode_pytorch.pyThe first run downloads the models (a few hundred MB), so give it a minute. Prefer smaller/faster models? Use:
python reference/spec_decode_pytorch.py --draft distilgpt2 --target gpt2You should see the generated text plus a speed stat:
Device: cuda
--- OUTPUT ---
The key idea behind speculative decoding is ...
--- STATS ---
avg tokens per target pass: 3.23 (>1 means speedup)
avg tokens per target pass: 3.23 means the big model ran once for every ~3
words instead of once per word — that's the speedup, live.
No install, no terminal. Just double-click
speculative_decoding_explainer.html
(or open it in any browser). It's fully self-contained. Drag the probability bars
and move the sliders to build intuition.
The real-model demo (option B) picks the best available accelerator for you — you never change the code:
| Your computer | What it uses automatically |
|---|---|
| NVIDIA GPU (Windows / Linux) | CUDA |
| Apple Silicon Mac (M1/M2/M3/M4) | Apple Metal (MPS) |
| No GPU / anything else | CPU (slower, but always works) |
Want to force one? Add --device cpu (or cuda, or mps). On Apple Silicon,
CPU-fallback for any unsupported operation is enabled automatically, so it "just
works."
All options for the real-model demo (spec_decode_pytorch.py):
| Option | Default | Meaning |
|---|---|---|
--draft |
distilgpt2 |
The small "guesser" model (Hugging Face name). |
--target |
gpt2-large |
The big "checker" model. |
--prompt |
"The key idea…" | The starting text. |
--gamma |
4 |
How many words to guess per round. Try 3–7. |
--max-new-tokens |
60 |
How many words to generate. |
--temperature |
1.0 |
Randomness. 0 = most likely word every time. |
--device |
auto |
auto, cuda, mps, or cpu. |
--seed |
0 |
Change for different random output. |
Example — a longer, more focused generation on the GPU:
python reference/spec_decode_pytorch.py \
--draft distilgpt2 --target gpt2 \
--prompt "Once upon a time" --gamma 5 \
--max-new-tokens 100 --temperature 0.7| Problem | Fix |
|---|---|
python: command not found |
Try python3 instead of python. On Windows, reinstall Python and tick "Add Python to PATH". |
No module named numpy |
Run pip install -r requirements.txt (and make sure your virtual env is activated). |
No module named torch |
That's only for option B — run pip install -r requirements-pytorch.txt. |
| Model download is slow / fails | It's a one-time download; retry on a stable connection. Use smaller models: --draft distilgpt2 --target gpt2. |
| Out of memory on GPU | Use smaller models (--target gpt2) or force CPU with --device cpu. |
avg tokens per target pass is low (~1) |
Normal if draft and target rarely agree. Use a draft from the same family as the target for higher agreement. |
pip installs to the wrong Python |
Use python -m pip install ... so pip matches the Python you're running. |
Confirms the two core claims — losslessness and the speedup math:
# with pytest:
pip install pytest
pytest
# or with no extra install:
python tests/test_losslessness.pyExpected: 6 passed.
New to the topic? Read the course in order (about an hour total):
guide/00-intuition.md— the whole idea, no mathguide/01-math.md— the accept/reject rule and why it's exactguide/02-algorithm.md— the step-by-step recipe + the codeguide/03-methods.md— the modern methods (token trees, Medusa, EAGLE)
Prefer PDFs? The same four chapters are in notes/ as designed,
print-ready documents (built from Typst source — rebuild with typst compile).
Notation used everywhere: p = the target (big/accurate) model · q = the
draft (small/fast) model · γ (gamma) = how many words the draft guesses
before the target checks.
MIT © 2026 Syed Azeez. The PDFs in papers/ belong to their respective
authors — see papers/README.md for arXiv links.