05-01 Query, Key, Value Intuition
Why this matters
In Attention Is All You Need, attention output is a weighted sum of value vectors using query-key compatibility weights.
Intuition first (no jargon)
Queries ask what information is needed, keys expose what each position offers, and values carry content to mix.
Paper grounding
- The paper defines attention as
Attention(Q, K, V) = softmax(QK^T / sqrt(d_k))V. - It describes the output as a weighted sum of values, where weights come from query-key compatibility.
Worked example
- Inputs:
xhasT = 3tokens anddModel = 4 - Shapes:
xis[3][4]Wq,Wk,Wvare[4][2]Q,K,Vbecome[3][2]
- One computed step:
Code walkthrough
jsexport function makeQKV(x, params) { // x: [T][dModel] // returns { Q, K, V } }
makeQKVprojects each token representation into separate query, key, and value vectors.
Your task
Implement makeQKV(x, params).
- Compute
Q = xWq + bq,K = xWk + bk,V = xWv + bv. - Keep dimensions consistent for all tokens.
- Return new arrays without mutating input.
Common mistakes
- Reused projection matrix for
Q,K, andV - Mixed
dModelanddKeyin output allocation - Missing shape assertions for
[T][dKey]and[T][dValue]
Precision note
Scaled dot-product attention uses query-key scores, 1 / sqrt(dk) scaling, and row-wise softmax.
Hints
- Reuse linear helpers from earlier units.
- Assert shapes in debug mode.
- Keep parameter names explicit.
Check your thinking
- Why are Q and K compared?
- Why is V separate from K?
- What would happen if Q, K, and V were identical always?
Stretch (optional)
Support optional shared projection for K and V and compare behavior.
Likely test focus
- Correct output shapes.
- Deterministic numeric outputs for fixed params.
- No input mutation.
What should improve
Your Q/K/V projections are now ready for scaled dot-product score computation.