Lesson 01-04

Unigram Baseline Generator

15 min
2 exports
4 tests

Lesson blocked by prerequisites

Complete and save a passing attempt for your active lesson before running this one.

Go to active lesson

Lesson workspace sections

Submission + Results
Not run yet

Seed 104 • Runtime includes 3 prerequisite modules
Test Results

Run tests to see case-by-case feedback.

Attempts0 saved

No saved attempts yet.

Lesson README

01-04 Unigram Baseline Generator

Why this matters

A unigram baseline closes the full loop from training counts to generated output.

Intuition first (no jargon)

If selection depends only on global frequency, output reflects corpus-wide token prevalence.

Code walkthrough

js
export function trainUnigram(text) {
  return { items: [], probs: [] };
}

export function generateUnigram(model, steps, rng = Math.random) {
  return "";
}

Your task

Implement trainUnigram and generateUnigram.

  • Convert counts into probabilities that sum to 1.
  • Store token list and probability list.
  • Generate steps tokens by repeated weighted sampling.

Hints

  • Reuse countChars and weightedRandom.
  • Protect against empty training text.
  • Keep output deterministic when seeded RNG is provided.

Check your thinking

  1. Why does unigram output ignore context?
  2. What kinds of patterns can this model never learn?
  3. Why is this still a useful milestone?

Stretch (optional)

Add a helper that reports top 5 most probable tokens.

Likely test focus

  • Probabilities are normalized.
  • Output length matches steps.
  • Handles empty/invalid inputs safely.

What should improve

You now have a complete baseline model for comparison with stronger context-aware models.

Bridge to next lesson

Next: represent text as token IDs for bigram learning.

Monaco Editor

Matches starter

Files

Editor is deferred on smaller screens to keep startup fast.

Autosave is enabled in local storage for this lesson.