Karpathy's pure-Python microGPT as a Strands Agents model provider — zero dependencies, pure autograd, trained from scratch.
Based on @karpathy's atomic GPT gist: "The most atomic way to train and run inference for a GPT in pure, dependency-free Python. This file is the complete algorithm. Everything else is just efficiency."
A complete GPT implementation — autograd engine, transformer architecture, tokenizer, Adam optimizer, training loop, and inference — in pure Python with zero dependencies. No PyTorch, no NumPy, no CUDA. Just math, random, and the algorithm.
Packaged as a proper Strands model provider so you can:
- Use it as a Model — drop-in Strands
Modelinterface for character-level generation - Use it as a Tool — train and generate from any Strands agent via tool calls
- Learn from it — the entire algorithm is readable, hackable, and documented
pip install strands-microgptRequirements: Python ≥3.10,
strands-agents. That's it. No GPU needed.
from strands_microgpt import MicroGPT
# Load dataset, build tokenizer, create model
model, tokenizer, docs = MicroGPT.from_dataset()
# Train (1000 steps on names.txt)
model.train_on_docs(docs, tokenizer, num_steps=1000)
# Generate new names
for name in model.generate(tokenizer, num_samples=10):
print(name)from strands import Agent
from strands_microgpt import MicroGPTModel
model = MicroGPTModel(num_steps=1000, temperature=0.5)
agent = Agent(model=model)
agent("Generate some names")from strands import Agent
from strands_microgpt import microgpt_train, microgpt_generate
# Use with Bedrock, OpenAI, or any model
agent = Agent(tools=[microgpt_train, microgpt_generate])
agent("Train a GPT on the names dataset for 500 steps, then generate 10 names")The complete algorithm in ~300 lines:
strands_microgpt/
├── engine.py # Value (autograd), Tokenizer, MicroGPT (transformer)
├── microgpt_model.py # Strands Model interface
└── tools/
├── microgpt_train.py # Training tool
└── microgpt_generate.py # Generation tool
from strands_microgpt import Value
a = Value(2.0)
b = Value(3.0)
c = a * b + a # builds computation graph
c.backward() # backpropagate gradients
print(a.grad) # 4.0 (dc/da = b + 1)Supports: +, *, -, /, **, relu(), exp(), log(), backward()
GPT-2 architecture with:
- RMSNorm (instead of LayerNorm)
- No biases
- ReLU (instead of GeLU)
- Multi-head causal attention
- Adam optimizer with linear LR decay
| Config | Default | Description |
|---|---|---|
n_layer |
1 | Transformer depth |
n_embd |
16 | Embedding dimension |
block_size |
16 | Context window |
n_head |
4 | Attention heads |
num_steps |
1000 | Training steps |
learning_rate |
0.01 | Initial LR (linear decay) |
temperature |
0.5 | Generation temperature |
Train on anything — names, poems, molecules, DNA, code:
from strands_microgpt import MicroGPT, Tokenizer
docs = ["the cat sat on the mat", "the dog sat on the log"] * 100
tokenizer = Tokenizer.from_docs(docs)
model = MicroGPT(vocab_size=tokenizer.vocab_size, n_embd=32, block_size=32)
model.train_on_docs(docs, tokenizer, num_steps=2000)
samples = model.generate(tokenizer, num_samples=10, temperature=0.7)# Save
model.save_checkpoint("model.json", tokenizer)
# Load
model, tokenizer, metadata = MicroGPT.load_checkpoint("model.json")
samples = model.generate(tokenizer, num_samples=10)| Example | Description |
|---|---|
| 01_basic_training.py | Train on names, generate new ones |
| 02_strands_agent.py | Use as a Strands Model provider |
| 03_tool_usage.py | Train/generate via tool calls |
| 04_custom_dataset.py | Train on custom text data |
| 05_autograd_exploration.py | Explore the autograd engine |
"Everything else is just efficiency." — @karpathy
This package exists to show that:
- A GPT is just math. No magic, no black boxes. The entire algorithm fits in your head.
- Strands Model interface is universal. If it can generate tokens, it can be a Strands model.
- Understanding > Using. Train a transformer from scratch to truly grok what LLMs do.
The model is tiny and slow (pure Python, no vectorization). For production, use Bedrock, OpenAI, or any real provider. For learning, this is the best code to read.
MicroGPT(vocab_size, n_layer=1, n_embd=16, block_size=16, n_head=4, seed=42).train_on_docs(docs, tokenizer, num_steps, learning_rate, log_every, callback)→List[float].generate(tokenizer, num_samples, temperature, max_length)→List[str].save_checkpoint(path, tokenizer, metadata)→ None.load_checkpoint(path)→(MicroGPT, Tokenizer, Dict)(classmethod).from_dataset(dataset_url, dataset_path, **kwargs)→(MicroGPT, Tokenizer, docs)(classmethod)
MicroGPTModel(dataset_url, num_steps=1000, temperature=0.5, ...)Drop-in replacement for any Strands model. Trains on first use, then generates.
Value(data) # scalar autograd nodeSupports: +, *, -, /, **, .relu(), .exp(), .log(), .backward()
microgpt_train(dataset_url, num_steps, n_layer, ...)— Train a modelmicrogpt_generate(checkpoint_path, num_samples, temperature)— Generate from checkpoint
- Karpathy's GPT gist — The original
- micrograd — Karpathy's autograd engine
- makemore — Character-level language modeling
- Strands Agents — The agent framework
- strands-cosmos — NVIDIA Cosmos VLM provider
MIT | Based on @karpathy's work | Built with Strands Agents