Lesson 06-01

Multi-Head Attention

12 min
3 exports
5 tests

Lesson blocked by prerequisites

Complete and save a passing attempt for your active lesson before running this one.

Go to active lesson

Lesson workspace sections

Submission + Results
Not run yet

Seed 601 • Runtime includes 20 prerequisite modules
Test Results

Run tests to see case-by-case feedback.

Attempts0 saved

No saved attempts yet.

Lesson README

06-01 Multi-Head Attention

Why this matters

Section 3.2.2 projects Q, K, and V into multiple subspaces so different heads can attend to different patterns.

Intuition first (no jargon)

Multiple smaller attentions run in parallel, then concatenate and project back to model width.

Paper grounding

  • The paper defines head_i = Attention(QW_i^Q, KW_i^K, VW_i^V).
  • Multi-head output is Concat(head_1, ..., head_h)W^O.

In the paper, queries, keys, and values are projected h times with learned linear projections before per-head attention.

Code walkthrough

js
export function splitHeads(x, numHeads) {}
export function combineHeads(heads) {}
export function multiHeadAttention(x, params, mask) {}

Your task

Implement head split/combine and full multi-head attention.

  • Reshape [T][dModel] into [H][T][dHead].
  • Run scaled masked attention per head.
  • Concatenate heads and project back to dModel.

Hints

  • Validate dModel % numHeads === 0.
  • Keep head dimension naming consistent.
  • Test split then combine round-trip.

Check your thinking

  1. Why can heads specialize?
  2. What does output projection do after concatenation?
  3. Why is shape bookkeeping critical here?

Stretch (optional)

Return per-head attention maps for visualization.

Likely test focus

  • Shape correctness at each stage.
  • Split/combine round-trip integrity.
  • Deterministic output with fixed params.

What should improve

Your implementation now supports split-attend-concatenate multi-head computation.

Bridge to next lesson

Next lesson: add token-wise feed-forward transformation.

Monaco Editor

Matches starter

Files

Editor is deferred on smaller screens to keep startup fast.

Autosave is enabled in local storage for this lesson.