Lesson 06-02

Feed-Forward Layer

10 min
2 exports
5 tests

Lesson blocked by prerequisites

Complete and save a passing attempt for your active lesson before running this one.

Go to active lesson

Lesson workspace sections

Submission + Results
Not run yet

Seed 602 • Runtime includes 21 prerequisite modules
Test Results

Run tests to see case-by-case feedback.

Attempts0 saved

No saved attempts yet.

Lesson README

06-02 Feed-Forward Layer

Why this matters

The Transformer applies a position-wise FFN: max(0, xW1 + b1)W2 + b2 after attention mixing.

Intuition first (no jargon)

After cross-position mixing, each position gets a deeper non-linear transform independently.

Paper grounding

  • The feed-forward sublayer is FFN(x) = max(0, xW_1 + b_1)W_2 + b_2.
  • The same FFN parameters are applied to each position independently.

Code walkthrough

js
export function ffnToken(x, params) {}
export function ffn(xSeq, params) {}

Your task

Implement FFN for one token and full sequence.

  • Compute xW1 + b1, apply activation.
  • Compute second projection back to dModel.
  • Process every token independently.

Hints

  • Reuse existing linear and activation helpers.
  • Keep xSeq.length unchanged.
  • Check hidden dimension shape carefully.

Check your thinking

  1. Why is FFN applied per token independently?
  2. What role does hidden expansion play?
  3. Why pair attention and FFN together?

Stretch (optional)

Try GELU activation and compare training smoothness later.

Likely test focus

  • Correct output shape [T][dModel].
  • Deterministic transform on toy inputs.
  • No cross-token mixing inside FFN.

What should improve

You can now implement the position-wise FFN stage used in Transformer blocks.

Bridge to next lesson

Next you add residual connections and normalization for stable deeper stacks.

Monaco Editor

Matches starter

Files

Editor is deferred on smaller screens to keep startup fast.

Autosave is enabled in local storage for this lesson.