02-03 Bigram Probabilities and Predict
Why this matters
Normalization turns raw transition counts into valid next-token probability distributions.
Intuition first (no jargon)
Each row must sum to 1 before it can be used for probabilistic prediction.
Code walkthrough
jsexport function bigramProbs(counts) {} export function predictNext(currentId, probs, rng = Math.random) {}
Your task
Implement bigramProbs and predictNext.
- Normalize each row so row sum is 1 when row has data.
- Define a fallback for zero rows (uniform or configured default).
- Sample next token from the chosen row.
Hints
- Compute row sum once per row.
- Reuse
weightedRandomfor sampling. - Keep fallback behavior explicit and testable.
Check your thinking
- Why are zero rows common on tiny datasets?
- Why must each row sum to 1?
- What tradeoff comes with uniform fallback?
Stretch (optional)
Add additive smoothing to reduce zero-probability outputs.
Likely test focus
- Row normalization correctness.
- Valid next-token ID range.
- Predict works on sparse rows.
What should improve
You can now query meaningful next-token probabilities from bigram statistics.
Bridge to next lesson
Next lesson: chain predictions to generate sequences.