kiriz/llm-basics

Step-by-step animated visualization of how an LLM turns a prompt into the next token. HuggingFace transformers under the hood; renders as standalone HTML.

★ 0Forks 0PythonGitHub ↗Compare

Project website ↗

educationhuggingfacellmmachine-learningnlptokenizationtransformersvisualization

README

llm-basics

Step-by-step animated visualization of how a language model turns a prompt into the next token. Built on HuggingFace transformers (no inference servers, no abstractions hidden) so every layer, hidden state, attention pattern, logit, and probability is captured and renderable as standalone HTML.

The point: demystify "the model is just a function over learned weights". Each step in the pipeline shows the actual numbers the model is operating on, for the actual prompt you ran.

Live demos

Hosted on GitHub Pages: https://kiriz.github.io/llm-basics/

How LLMs are wired (explainer)

A written page for engineers who think in APIs: the two stacked layers (an ordinary REST contract on top, int[] → float[vocab] underneath), why the model is stateless, and why prompt injection has no prepared-statement equivalent.

Its claims about prompt format are shown rather than asserted. The page renders one identical pair of messages through six model families' own published chat_template files — distilgpt2, TinyLlama, Qwen2.5, NVIDIA's Llama-derived Nemotron, Mistral and OpenAI's gpt-oss — each linked to the source file. Three things fall out that prose alone would not settle:

  • distilgpt2 has no chat template at all — roles are taught in post-training, not built into the architecture.
  • Mistral has no system role; it folds system text into the first user turn. gpt-oss renames system to developer and prepends its own synthesised system block.
  • Nemotron injects <AVAILABLE_TOOLS>[] even when no tools were declared.

Regenerate the exhibits with python fetch-chat-templates.py (fetches tokenizer configs only — no weights). Meta's Llama and Google's Gemma are licence-gated, so the page cites Meta's published prompt-format docs rather than showing an artefact.

What the demo shows

Each animated demo walks through 9 steps:

  1. Input — the raw prompt string
  2. Tokenize — text → integer ids (BPE / sentencepiece)
  3. Embed — id → 768/2048-dim vector via the embedding table
  4. Layers — residual-stream evolution through stacked transformer blocks
  5. Logits — final hidden state projected by the LM head into one score per vocab token
  6. Softmax — logits → probabilities
  7. Sample — pick one token (argmax for greedy)
  8. Loop — repeat until EOS or max tokens
  9. Done — final output text

A right-side "where am I in the stack?" panel shows whether each step happens in the frontend, inference runtime (HuggingFace transformers), or model (the weights themselves).

Highlights

Step 4 — residual stream evolution through layers. Each row is the same last-token vector after one more transformer block. The norm bar on the right shows how much "magnitude" each layer adds.

Residual stream heatmap across layers

Step 5 — LM head as similarity search. The final hidden state gets dot-producted against every vocab token's W-row in parallel; top-3 candidates shown with their actual W-row vectors and logit scores.

LM head similarity search

Step 8 — the autoregressive loop. 138 generated tokens (TinyLlama answering "Name three primary colors.") with the prompt in muted gray and EOS as a red chip at the end.

Loop with 138 generated tokens

Step 8, expanded — the matmul that emits EOS. At step 138 the model picked </s> (logit +16.59) over They (+15.80) and The (+13.40). Click any candidate row to see all 2048 dims of the W-row that made it the winner.

EOS-step LM head matmul

Inside one transformer block

A separate 6-substep slideshow opens up the box at step 4 — one chosen (layer, head), all the actual numbers, animated. Live: distilgpt2 · TinyLlama.

Step 1 — LayerNorm + Q/K/V projection. The block input vector gets normalized, then split into three vectors of width head_dim for the chosen head.

Q K V projection

Step 2 — scores Q · Kᵀ. Every (query, key) pair gets a dot product. Cells fill row-by-row in a stagger animation; upper-triangle is masked (causal). 20 tokens × 20 keys = 400 dot products, only 210 unmasked.

Q K^T scores matrix

Step 3 — softmax morph. The scores row for the spotlighted query token morphs through three stages: raw scores (some negative) → e^x (all positive) → divide by Σ (probabilities sum to 1).

Softmax morph final stage

Step 5 — FFN expand + activation. The post-attention residual goes through LN2, then a linear up_proj from hidden_size to ffn_dim (here 2048 → 5632, 2.75× wider). Each cell passes through the activation curve (SiLU for Llama, GELU for GPT-2).

FFN expand to 5632 dims

Run it locally

Requires uv (https://astral.sh/uv) and a working Python 3.11+.

./setup.sh                                          # install deps + warm up distilgpt2
./try-prompt.sh "Apple is round and"                # run any prompt through distilgpt2
./try-prompt.sh "Hi" --model TinyLlama/TinyLlama-1.1B-Chat-v1.0 --max 60

Output HTMLs go to ./out/. Re-run with the same prompt + model: cache hit, no model load.

Layout

src/llm_trace/         # the package
  cli.py               # typer CLI: run | render | list-cache | clear-cache
  collector.py         # loads model, runs forward pass, captures intermediates
  cache.py             # disk cache (.npz arrays + .json sidecar, atomic writes)
  trace_data.py        # TraceData dataclass — torch-free, the spine of the data model
  config.py            # YAML config loader
  embeddings.py        # standalone embedding-space PCA explorer
  renderers/           # pure TraceData → output (no torch imports allowed)
    terminal.py        # rich-formatted terminal output
    png.py             # matplotlib summary plot
    html.py            # static HTML renderer
    animated_v3.py     # the slideshow demo (the highlighted artifact)

tests/llm_trace/       # unit + integration tests
docs/                  # GitHub Pages content (the rendered demos)
trace.yaml             # default config (models, prompts, generation params)
setup.sh               # one-time install via uv
try-prompt.sh          # one-shot prompt → animated HTML

Why HuggingFace transformers

We bypass higher-throughput runtimes (vLLM, TGI, llama.cpp) on purpose. They optimize for serving and hide intermediates. The transformers Python forward pass exposes output_hidden_states and output_attentions directly — that's what makes step-by-step visualization possible.

License

MIT.

Contributors

kiriz

Issues