05-02 Attention Scores and Weights
Why this matters
Scaled dot-product attention computes QK^T / sqrt(dKey) and applies row-wise softmax to produce attention weights.
Intuition first (no jargon)
Each query row scores all keys, then softmax turns scores into a normalized weighting distribution.
Paper grounding
- Scaled dot-product attention divides by
sqrt(d_k)before softmax. - The paper explains this scaling prevents large dot-product magnitudes that would push softmax into very small-gradient regions.
Worked example
- Inputs:
Q = [[1, 0], [0, 1]]K = [[1, 1], [1, -1]]T = 2,dKey = 2
- Shapes:
Q [2][2],K [2][2]S = QK^Tis[2][2]weightsafter row softmax is[2][2]
- One computed step:
S = [[1, 1], [1, -1]]- Scale scores by
1 / sqrt(dKey)before softmax. - Row-0 softmax after scaling is approximately
[0.50, 0.50]; row-1 is approximately[0.804, 0.196].
- Output: token 1 attends much more to position 0 than position 1.
Code walkthrough
jsexport function attentionScores(Q, K, scale = true) {} export function attentionWeights(scores) {}
attentionScorescomputesQK^Tand applies optional1 / sqrt(dKey)scaling.attentionWeightsnormalizes each query row independently.
Your task
Implement score and weight computation.
- Compute score matrix
S = Q * K^T. - If
scale, divide bysqrt(dKey). - Softmax each row to get weights.
Common mistakes
- Softmax over full matrix instead of per row
- Missing
sqrt(dKey)scaling before softmax
Precision note
Decoder self-attention: position i attends only to positions up to and including i.
Hints
- Reuse stable softmax from Unit 03.
- Row
icorresponds to tokeniasking for context. - Verify each weight row sums near 1.
Check your thinking
- Why apply softmax per row?
- Why use scaling by
sqrt(dKey)? - What does a sharp distribution mean?
Stretch (optional)
Add function to return top-2 attended positions per token.
Likely test focus
- Score matrix shape
[T][T]. - Row sums near 1 after softmax.
- Correct scaling behavior.
What should improve
You can now compute stable attention weight matrices that sum to 1 per query row.
Bridge to next lesson
Next lesson: apply weights to value vectors.