Lesson 02-02

Build Bigram Counts

12 min
1 export
3 tests

Lesson blocked by prerequisites

Complete and save a passing attempt for your active lesson before running this one.

Go to active lesson

Lesson workspace sections

Submission + Results
Not run yet

Seed 202 • Runtime includes 5 prerequisite modules
Test Results

Run tests to see case-by-case feedback.

Attempts0 saved

No saved attempts yet.

Lesson README

02-02 Build Bigram Counts

Why this matters

Bigram counts capture first-order local order by tracking which token follows which.

Intuition first (no jargon)

A transition matrix stores one-step sequence memory for each current token.

Code walkthrough

js
export function buildBigramCounts(ids, vocabSize) {
  // returns vocabSize x vocabSize matrix of counts
}

Your task

Implement buildBigramCounts(ids, vocabSize).

  • Create a zero-initialized 2D array.
  • For each adjacent pair (ids[i], ids[i + 1]), increment count.
  • Skip counting when sequence length is less than 2.

Hints

  • Row index is current token.
  • Column index is next token.
  • Validate ID range before incrementing.

Check your thinking

  1. What does counts[a][b] mean?
  2. Why do we need both row and column dimensions?
  3. How is this different from unigram counts?

Stretch (optional)

Track sentence-start and sentence-end markers as special IDs.

Likely test focus

  • Correct matrix shape.
  • Correct increments for toy ID sequences.
  • Handles edge lengths safely.

What should improve

Your model now represents local sequential structure beyond unigram frequency.

Bridge to next lesson

Next: convert count rows into probabilities.

Monaco Editor

Matches starter

Files

Editor is deferred on smaller screens to keep startup fast.

Autosave is enabled in local storage for this lesson.