05-04 Causal Mask (No Peeking)
Why this matters
Decoder self-attention must block future-token access so position i attends only to positions <= i.
Intuition first (no jargon)
Apply the mask before softmax so illegal future positions receive effectively zero probability.
Paper grounding
- For decoder self-attention, illegal future positions are masked before softmax by adding a large negative value to those logits.
- This enforces that each position attends only to itself and earlier positions.
Worked example
- Inputs: sequence length
T = 4 - Shapes:
- mask is
[4][4] - scores are
[4][4] - masked scores remain
[4][4]
- mask is
- One computed step:
- valid causal mask rows are:
- row 0:
[1, 0, 0, 0] - row 1:
[1, 1, 0, 0] - row 2:
[1, 1, 1, 0] - row 3:
[1, 1, 1, 1]
- row 0:
- row 2 cannot assign attention to position 3.
- valid causal mask rows are:
Code walkthrough
jsexport function causalMask(T) {} export function applyMask(scores, mask) {}
causalMask: lower-triangular allow matrixapplyMask: large negative value for disallowed score cells
Your task
Implement masking utilities and integrate into attention.
causalMask(T)allows positionsj <= ionly.- Replace masked score entries with a very negative number.
- Confirm masked weights for future positions become near zero.
Common mistakes
- Flipped mask orientation (
j > iallowed) - Mask applied after softmax instead of before
- Non-negative sentinel value used for masked scores
Precision note
This curriculum uses -1e9 as a practical stand-in for negative infinity.
Hints
- Use
-1e9as a practical negative infinity. - Apply mask before softmax.
- Test with small
T = 4examples.
Check your thinking
- Why does masking happen on scores, not weights?
- What bug appears if mask is reversed?
- Why is this needed in both train and generate modes?
Stretch (optional)
Support an additional padding mask for variable-length batches.
Likely test focus
- Correct lower-triangular mask structure.
- Future weights near zero.
- Causal constraints preserved across rows.
What should improve
Your attention pipeline now enforces autoregressive constraints correctly.
Bridge to next lesson
Next: run multiple attention heads in parallel.