Curriculum
Build Your Own LLM
Choose a lesson, open the coding workspace, and pass tests as you build a tiny LLM from zero.
Unit lessons
Unit 01: First Wins (Coding Basics)
4 bite-sized lessons · 0/4 completed
| Lesson | Title | Status | Best Score | Duration | Action |
|---|---|---|---|---|---|
| 01-01 | First Green Test Text normalization prevents artificial vocabulary inflation from casing and spacing artifacts, so downstream counts represent real language patterns. Prerequisites: None | Not started | - | 10 min | Open lesson |
| 01-02 | Counting What the Model Sees Token-frequency counting is the first measurable signal a language model can extract from data. Prerequisites: 01-01 | Not started | - | 10 min | Complete 01-01 |
| 01-03 | Random Choice for Generation Weighted sampling turns static token probabilities into diverse generation behavior. Prerequisites: 01-02 | Not started | - | 12 min | Complete 01-02 |
| 01-04 | Unigram Baseline Generator A unigram baseline closes the full loop from training counts to generated output. Prerequisites: 01-03 first-text-generator | Not started | - | 15 min | Complete 01-03 |
Unit 02: First Working LM (Bigram)
4 bite-sized lessons · 0/4 completed
| Lesson | Title | Status | Best Score | Duration | Action |
|---|---|---|---|---|---|
| 02-01 | Tokenize and Vocabulary Vocabulary mapping is the numeric interface between text and model computation. Prerequisites: 01-04 | Not started | - | 12 min | Complete 01-04 |
| 02-02 | Build Bigram Counts Bigram counts capture first-order local order by tracking which token follows which. Prerequisites: 02-01 | Not started | - | 12 min | Complete 02-01 |
| 02-03 | Bigram Probabilities and Predict Normalization turns raw transition counts into valid next-token probability distributions. Prerequisites: 02-02 | Not started | - | 12 min | Complete 02-02 |
| 02-04 | Generate with Bigram Bigram generation is the first model in this curriculum that conditions on immediate context. Prerequisites: 02-03 first-working-learned-lm | Not started | - | 15 min | Complete 02-03 |
Unit 03: Trainable Neural LM
4 bite-sized lessons · 0/4 completed
| Lesson | Title | Status | Best Score | Duration | Action |
|---|---|---|---|---|---|
| 03-01 | Make Context Target Examples Supervised next-token training requires explicit (context, target) pairs from token streams. Prerequisites: 02-04 | Not started | - | 12 min | Complete 02-04 |
| 03-02 | One-Hot and Linear Logits A linear logit layer is the simplest trainable scorer from token representation to vocabulary scores. Prerequisites: 03-01 | Not started | - | 12 min | Complete 03-01 |
| 03-03 | Softmax and Cross-Entropy Transformer training optimizes next-token probabilities via softmax outputs and cross-entropy loss. Prerequisites: 03-02 | Not started | - | 12 min | Complete 03-02 |
| 03-04 | Gradient Step That Reduces Loss Gradient updates are the mechanism that turns prediction errors into parameter improvements. Prerequisites: 03-03 | Not started | - | 15 min | Complete 03-03 |
Unit 04: Embeddings and Better Context
4 bite-sized lessons · 0/4 completed
| Lesson | Title | Status | Best Score | Duration | Action |
|---|---|---|---|---|---|
| 04-01 | Embedding Lookup Embedding lookup replaces sparse one-hot inputs with dense trainable token representations. Prerequisites: 03-04 | Not started | - | 12 min | Complete 03-04 |
| 04-02 | Combine Context Vectors A language head needs one context representation per example before scoring next-token logits. Prerequisites: 04-01 | Not started | - | 10 min | Complete 04-01 |
| 04-03 | MLP Language Head An MLP head adds non-linearity so the predictor can model richer token interactions. Prerequisites: 04-02 | Not started | - | 12 min | Complete 04-02 |
| 04-04 | Train Neural Language Model This lesson connects embeddings, context composition, logits, loss, and updates in one reproducible loop. Prerequisites: 04-03 neural-beats-bigram | Not started | - | 15 min | Complete 04-03 |
Unit 05: Attention Fundamentals
4 bite-sized lessons · 0/4 completed
| Lesson | Title | Status | Best Score | Duration | Action |
|---|---|---|---|---|---|
| 05-01 | Query, Key, Value Intuition In Attention Is All You Need, attention output is a weighted sum of value vectors using query-key compatibility weights. Prerequisites: 04-04 | Not started | - | 12 min | Complete 04-04 |
| 05-02 | Attention Scores and Weights Scaled dot-product attention computes QK^T / sqrt(dKey) and applies row-wise softmax to produce attention weights. Prerequisites: 05-01 | Not started | - | 12 min | Complete 05-01 |
| 05-03 | Weighted Sum of Values The paper's attention output multiplies weights by value vectors to produce contextualized representations. Prerequisites: 05-02 | Not started | - | 10 min | Complete 05-02 |
| 05-04 | Causal Mask (No Peeking) Decoder self-attention must block future-token access so position i attends only to positions <= i. Prerequisites: 05-03 | Not started | - | 12 min | Complete 05-03 |
Unit 06: Transformer Block
4 bite-sized lessons · 0/4 completed
| Lesson | Title | Status | Best Score | Duration | Action |
|---|---|---|---|---|---|
| 06-01 | Multi-Head Attention Section 3.2.2 projects Q, K, and V into multiple subspaces so different heads can attend to different patterns. Prerequisites: 05-04 | Not started | - | 12 min | Complete 05-04 |
| 06-02 | Feed-Forward Layer The Transformer applies a position-wise FFN: max(0, xW1 + b1)W2 + b2 after attention mixing. Prerequisites: 06-01 | Not started | - | 10 min | Complete 06-01 |
| 06-03 | Residual and LayerNorm Residual paths preserve signal, and normalization stabilizes scale across stacked sublayers. Prerequisites: 06-02 | Not started | - | 12 min | Complete 06-02 |
| 06-04 | Full Transformer Block Forward Pass A full block composes attention, residual/normalization, FFN, and residual/normalization in a fixed order. Prerequisites: 06-03 full-transformer-block | Not started | - | 15 min | Complete 06-03 |
Unit 07: Tiny GPT Training Loop
4 bite-sized lessons · 0/4 completed
| Lesson | Title | Status | Best Score | Duration | Action |
|---|---|---|---|---|---|
| 07-01 | Positional Information Section 3.5 adds positional information because attention alone has no recurrence or convolutional order signal. Prerequisites: 06-04 | Not started | - | 12 min | Complete 06-04 |
| 07-02 | Stack Blocks Into MiniGPT Stacking blocks creates the full forward path from embeddings to vocabulary logits. Prerequisites: 07-01 | Not started | - | 15 min | Complete 07-01 |
| 07-03 | Batching and Train Step Batching makes optimization efficient and keeps tensor contracts consistent across updates. Prerequisites: 07-02 | Not started | - | 15 min | Complete 07-02 |
| 07-04 | Validation and Checkpoints Validation estimates generalization, and checkpoints preserve exact model state for reproducible comparison. Prerequisites: 07-03 | Not started | - | 12 min | Complete 07-03 |
Unit 08: Capstone Quality and Interpretability
4 bite-sized lessons · 0/4 completed
| Lesson | Title | Status | Best Score | Duration | Action |
|---|---|---|---|---|---|
| 08-01 | Temperature Sampling Temperature rescales logits during sampling to control distribution sharpness without retraining weights. Prerequisites: 07-04 | Not started | - | 10 min | Complete 07-04 |
| 08-02 | Top-k Sampling Top-k sampling restricts decoding to the highest-scoring candidates before drawing a token. Prerequisites: 08-01 | Not started | - | 10 min | Complete 08-01 |
| 08-03 | Visualize Attention Maps Attention-map export makes per-head weighting behavior observable and debuggable. Prerequisites: 08-02 | Not started | - | 12 min | Complete 08-02 |
| 08-04 | Final Model Comparison A fair capstone comparison keeps prompt, seed, and evaluation protocol fixed across model variants. Prerequisites: 08-03 tiny-gpt-capstone | Not started | - | 15 min | Complete 08-03 |