Lesson 08-02

Top-k Sampling

10 min
1 export
4 tests

Lesson blocked by prerequisites

Complete and save a passing attempt for your active lesson before running this one.

Go to active lesson

Lesson workspace sections

Submission + Results
Not run yet

Seed 802 • Runtime includes 29 prerequisite modules
Test Results

Run tests to see case-by-case feedback.

Attempts0 saved

No saved attempts yet.

Lesson README

08-02 Top-k Sampling

Why this matters

Top-k sampling restricts decoding to the highest-scoring candidates before drawing a token.

Intuition first (no jargon)

Filter to the top k logits first, then sample within that reduced candidate set.

Paper grounding

  • The paper evaluates translation with beam search (beam size 4) and length normalization.
  • Top-k here is an explicit curriculum extension used to study decoding trade-offs.

Code walkthrough

js
export function sampleTopK(logits, k, temperature = 1, rng = Math.random) {}

Your task

Implement top-k filtered sampling.

  • Select indices of highest k logits.
  • Apply temperature and softmax only to those indices.
  • Return sampled ID from that subset.

Hints

  • Clamp k to [1, vocabSize].
  • Keep mapping between filtered and original IDs.
  • Test with k = 1 as deterministic argmax.

Check your thinking

  1. How does restricting candidates change output style?
  2. What is lost when k is too small?
  3. Why still keep temperature after top-k?

Stretch (optional)

Implement top-p (nucleus) sampling and compare outputs.

Likely test focus

  • Output always from top-k set.
  • Correct behavior for k = 1 and k >= vocab.
  • Stable handling of edge cases.

What should improve

You can now enforce candidate filtering before sampling for tighter generation control.

Bridge to next lesson

Next lesson: export attention maps.

Monaco Editor

Matches starter

Files

Editor is deferred on smaller screens to keep startup fast.

Autosave is enabled in local storage for this lesson.