Lesson 02-04

Generate with Bigram

15 min
2 exports
4 tests

Lesson blocked by prerequisites

Complete and save a passing attempt for your active lesson before running this one.

Go to active lesson

Lesson workspace sections

Submission + Results
Not run yet

Seed 204 • Runtime includes 7 prerequisite modules
Test Results

Run tests to see case-by-case feedback.

Attempts0 saved

No saved attempts yet.

Lesson README

02-04 Generate with Bigram

Why this matters

Bigram generation is the first model in this curriculum that conditions on immediate context.

Intuition first (no jargon)

The current token selects the next-token distribution at each generation step.

Code walkthrough

js
export function generateBigram(startId, steps, probs, rng = Math.random) {}

export function averageSurprise(ids, probs) {}

Your task

Implement generation and a simple quality metric.

  • Generate IDs autoregressively from startId.
  • Decode output to text.
  • Compute average surprise from predicted probabilities on a sequence.

Hints

  • Surprise is -log(p) averaged over steps.
  • Guard against p = 0 with a tiny epsilon.
  • Compare bigram and unigram on the same validation text.

Check your thinking

  1. Why is lower average surprise better?
  2. Why can bigram still fail on long dependencies?
  3. What does autoregressive mean in your own words?

Stretch (optional)

Add a helper that prints side-by-side unigram and bigram samples.

Likely test focus

  • Generated length and ID validity.
  • Surprise value is finite.
  • Bigram beats unigram on toy validation data.

What should improve

You can now generate context-sensitive samples and compare against unigram output.

Bridge to next lesson

Next lesson: move from count tables to trainable neural parameters.

Monaco Editor

Matches starter

Files

Editor is deferred on smaller screens to keep startup fast.

Autosave is enabled in local storage for this lesson.