Lesson 04-03

MLP Language Head

12 min
2 exports
5 tests

Lesson blocked by prerequisites

Complete and save a passing attempt for your active lesson before running this one.

Go to active lesson

Lesson workspace sections

Submission + Results
Not run yet

Seed 403 • Runtime includes 14 prerequisite modules
Test Results

Run tests to see case-by-case feedback.

Attempts0 saved

No saved attempts yet.

Lesson README

04-03 MLP Language Head

Why this matters

An MLP head adds non-linearity so the predictor can model richer token interactions.

Intuition first (no jargon)

Stacked linear layers with activation can represent patterns a single linear map cannot.

Code walkthrough

js
export function relu(x) {}
export function mlpForward(h, params) {}

Your task

Implement relu and mlpForward.

  • First affine transform: h1 = hW1 + b1.
  • Apply ReLU elementwise.
  • Second affine transform to vocab logits.

Hints

  • ReLU is Math.max(0, x).
  • Keep matrix dimensions documented.
  • Reuse your linear helper from earlier lessons.

Check your thinking

  1. What pattern can ReLU model that linear cannot?
  2. Why can dead ReLU units happen?
  3. Why do we still output logits at the end?

Stretch (optional)

Add tanh as an alternate activation and compare loss trend.

Likely test focus

  • Correct forward output shape.
  • Correct ReLU behavior on negatives.
  • Deterministic logits for fixed params.

What should improve

Your head can now model non-linear relationships in local context representations.

Bridge to next lesson

Now combine embeddings, context, and MLP into a full neural language model training loop.

Monaco Editor

Matches starter

Files

Editor is deferred on smaller screens to keep startup fast.

Autosave is enabled in local storage for this lesson.