01-04 Unigram Baseline Generator
Why this matters
A unigram baseline closes the full loop from training counts to generated output.
Intuition first (no jargon)
If selection depends only on global frequency, output reflects corpus-wide token prevalence.
Code walkthrough
jsexport function trainUnigram(text) { return { items: [], probs: [] }; } export function generateUnigram(model, steps, rng = Math.random) { return ""; }
Your task
Implement trainUnigram and generateUnigram.
- Convert counts into probabilities that sum to 1.
- Store token list and probability list.
- Generate
stepstokens by repeated weighted sampling.
Hints
- Reuse
countCharsandweightedRandom. - Protect against empty training text.
- Keep output deterministic when seeded RNG is provided.
Check your thinking
- Why does unigram output ignore context?
- What kinds of patterns can this model never learn?
- Why is this still a useful milestone?
Stretch (optional)
Add a helper that reports top 5 most probable tokens.
Likely test focus
- Probabilities are normalized.
- Output length matches
steps. - Handles empty/invalid inputs safely.
What should improve
You now have a complete baseline model for comparison with stronger context-aware models.
Bridge to next lesson
Next: represent text as token IDs for bigram learning.