06-04 Full Transformer Block Forward Pass
Why this matters
A full block composes attention, residual/normalization, FFN, and residual/normalization in a fixed order.
Intuition first (no jargon)
Correct sublayer ordering is what makes blocks stackable without shape or causality breaks.
Paper grounding
- Transformer layers combine self-attention and feed-forward sublayers with residual plus normalization wrappers.
- Decoder stacks preserve causal masking inside self-attention while keeping model width fixed across layers.
Code walkthrough
jsexport function transformerBlock(x, params, mask) { // attention path // residual + norm // ffn path // residual + norm }
Your task
Implement block composition in one function.
- Run masked multi-head attention.
- Apply residual and layer norm.
- Run FFN then another residual and layer norm.
- Return sequence with shape
[T][dModel].
Hints
- Follow one consistent order (pre-norm or post-norm).
- Keep variable names explicit to avoid wiring mistakes.
- Add quick shape assertions between stages.
Check your thinking
- Why does order matter in block composition?
- Where does causality enter this block?
- Why is this block stackable?
Stretch (optional)
Return intermediate tensors for debugging and visualization.
Likely test focus
- Correct forward order.
- Correct output shape.
- Respect for causal mask through attention stage.
What should improve
Your block forward pass now preserves shape and causal behavior end to end.
Bridge to next lesson
Next lesson: add positional signal and stack blocks.