08-03 Visualize Attention Maps
Why this matters
Attention-map export makes per-head weighting behavior observable and debuggable.
Intuition first (no jargon)
A head map is a T x T matrix showing how each query position weights each key position.
Paper grounding
- The paper includes qualitative attention analyses showing that different heads capture different dependency patterns.
- Inspecting per-head maps helps verify that masking and head separation behave as intended.
Code walkthrough
jsexport function extractAttentionMaps(cache) { // return serializable [layer][head][T][T] }
Your task
Implement attention map extraction utilities.
- Collect attention weights from each layer and head.
- Return JSON-serializable arrays for frontend rendering.
- Include token labels for axes.
Hints
- Keep extraction separate from rendering code.
- Validate each row sums near 1.
- Verify future positions are near zero under causal mask.
Check your thinking
- What pattern might punctuation show in attention?
- Why can different heads attend differently?
- Why should maps be captured during forward pass?
Stretch (optional)
Add summary metrics: diagonal strength and average entropy per head.
Likely test focus
- Correct nested shapes.
- Numeric normalization checks.
- Causal constraints preserved in extracted maps.
What should improve
You can now emit stable attention tensors for inspection and analysis.
Bridge to next lesson
Final lesson: compare model versions.