Step-by-step animated visualization of how a language model turns a prompt into the next token. Built on HuggingFace transformers (no inference servers, no abstractions hidden) so every layer, hidden state, attention pattern, logit, and probability is captured and renderable as standalone HTML.
The point: demystify "the model is just a function over learned weights". Each step in the pipeline shows the actual numbers the model is operating on, for the actual prompt you ran.
Hosted on GitHub Pages: https://kiriz.github.io/llm-basics/
- How LLMs are wired — the input/output contract — the explainer: an LLM as an API contract, with the real chat templates from six model families
- distilgpt2 — "January, February, March," — 6-layer GPT-2 base model completing a sequence
- TinyLlama — "Name three primary colors." — 22-layer chat-tuned Llama answering a question
- Inside one transformer block (distilgpt2) — 6-substep deep-dive: LN1 → Q/K/V → scores → softmax → FFN
- Inside one transformer block (TinyLlama) — same, on the larger Llama-family architecture (SwiGLU, GQA)
- Embedding-space scatter (distilgpt2) — PCA projection of all 50,257 vocab tokens
- Embedding-space scatter (TinyLlama) — same, for TinyLlama's 32,000 vocab
A written page for engineers who think in APIs: the two stacked layers (an ordinary REST
contract on top, int[] → float[vocab] underneath), why the model is stateless, and why
prompt injection has no prepared-statement equivalent.
Its claims about prompt format are shown rather than asserted. The page renders one
identical pair of messages through six model families' own published chat_template files —
distilgpt2, TinyLlama, Qwen2.5, NVIDIA's Llama-derived Nemotron, Mistral and OpenAI's
gpt-oss — each linked to the source file. Three things fall out that prose alone would not
settle:
- distilgpt2 has no chat template at all — roles are taught in post-training, not built into the architecture.
- Mistral has no system role; it folds system text into the first user turn. gpt-oss
renames
systemtodeveloperand prepends its own synthesised system block. - Nemotron injects
<AVAILABLE_TOOLS>[]even when no tools were declared.
Regenerate the exhibits with python fetch-chat-templates.py (fetches tokenizer configs
only — no weights). Meta's Llama and Google's Gemma are licence-gated, so the page cites
Meta's published prompt-format docs rather than showing an artefact.
Each animated demo walks through 9 steps:
- Input — the raw prompt string
- Tokenize — text → integer ids (BPE / sentencepiece)
- Embed — id → 768/2048-dim vector via the embedding table
- Layers — residual-stream evolution through stacked transformer blocks
- Logits — final hidden state projected by the LM head into one score per vocab token
- Softmax — logits → probabilities
- Sample — pick one token (argmax for greedy)
- Loop — repeat until EOS or max tokens
- Done — final output text
A right-side "where am I in the stack?" panel shows whether each step happens in the frontend, inference runtime (HuggingFace transformers), or model (the weights themselves).
Step 4 — residual stream evolution through layers. Each row is the same last-token vector after one more transformer block. The norm bar on the right shows how much "magnitude" each layer adds.
Step 5 — LM head as similarity search. The final hidden state gets dot-producted against every vocab token's W-row in parallel; top-3 candidates shown with their actual W-row vectors and logit scores.
Step 8 — the autoregressive loop. 138 generated tokens (TinyLlama answering "Name three primary colors.") with the prompt in muted gray and EOS as a red chip at the end.
Step 8, expanded — the matmul that emits EOS. At step 138 the model picked </s> (logit +16.59) over They (+15.80) and The (+13.40). Click any candidate row to see all 2048 dims of the W-row that made it the winner.
A separate 6-substep slideshow opens up the box at step 4 — one chosen (layer, head), all the actual numbers, animated. Live: distilgpt2 · TinyLlama.
Step 1 — LayerNorm + Q/K/V projection. The block input vector gets normalized, then split into three vectors of width head_dim for the chosen head.
Step 2 — scores Q · Kᵀ. Every (query, key) pair gets a dot product. Cells fill row-by-row in a stagger animation; upper-triangle is masked (causal). 20 tokens × 20 keys = 400 dot products, only 210 unmasked.
Step 3 — softmax morph. The scores row for the spotlighted query token morphs through three stages: raw scores (some negative) → e^x (all positive) → divide by Σ (probabilities sum to 1).
Step 5 — FFN expand + activation. The post-attention residual goes through LN2, then a linear up_proj from hidden_size to ffn_dim (here 2048 → 5632, 2.75× wider). Each cell passes through the activation curve (SiLU for Llama, GELU for GPT-2).
Requires uv (https://astral.sh/uv) and a working Python 3.11+.
./setup.sh # install deps + warm up distilgpt2
./try-prompt.sh "Apple is round and" # run any prompt through distilgpt2
./try-prompt.sh "Hi" --model TinyLlama/TinyLlama-1.1B-Chat-v1.0 --max 60Output HTMLs go to ./out/. Re-run with the same prompt + model: cache hit, no model load.
src/llm_trace/ # the package
cli.py # typer CLI: run | render | list-cache | clear-cache
collector.py # loads model, runs forward pass, captures intermediates
cache.py # disk cache (.npz arrays + .json sidecar, atomic writes)
trace_data.py # TraceData dataclass — torch-free, the spine of the data model
config.py # YAML config loader
embeddings.py # standalone embedding-space PCA explorer
renderers/ # pure TraceData → output (no torch imports allowed)
terminal.py # rich-formatted terminal output
png.py # matplotlib summary plot
html.py # static HTML renderer
animated_v3.py # the slideshow demo (the highlighted artifact)
tests/llm_trace/ # unit + integration tests
docs/ # GitHub Pages content (the rendered demos)
trace.yaml # default config (models, prompts, generation params)
setup.sh # one-time install via uv
try-prompt.sh # one-shot prompt → animated HTML
We bypass higher-throughput runtimes (vLLM, TGI, llama.cpp) on purpose. They optimize for serving and hide intermediates. The transformers Python forward pass exposes output_hidden_states and output_attentions directly — that's what makes step-by-step visualization possible.
MIT.







