04-04 Train Neural Language Model
Why this matters
This lesson connects embeddings, context composition, logits, loss, and updates in one reproducible loop.
Intuition first (no jargon)
A full train/eval cycle should produce logits, scalar loss, and trackable metric history.
Worked example
- Inputs: one batch with
B = 2,T = 4, vocab sizeV = 20 - Shapes:
- tokens
[B][T] -> embeddings [B][T][dModel] - model output logits
[B][T][V] - loss scalar
[]
- tokens
- One computed step:
- If validation loss for an epoch is
2.10, then surprise isexp(2.10) = 8.17
- If validation loss for an epoch is
- Output: report
lossand optional surprise.
Code walkthrough
jsexport function neuralForward(contextIds, params) {} export function trainNeuralLM(trainSet, valSet, params, cfg) {}
neuralForwardmaps context IDs to next-token logits.trainNeuralLMreturns train/validation metric history.
Your task
Implement training and evaluation flow.
- Build forward pass with embedding lookup and MLP head.
- Train for several epochs with configurable learning rate.
- Report train and validation loss each epoch.
- Optionally report surprise as
Math.exp(loss).
Common mistakes
- Changed seed or split between baseline comparisons
- Mixed context lengths across runs
Precision note
This lesson uses small dimensions so optimization behavior stays inspectable.
Hints
- Keep seed fixed when comparing baselines.
- Start with tiny model dimensions for fast runs.
- Early stopping can prevent overfitting on tiny data.
Check your thinking
- Why can train loss drop while validation loss rises?
- Why is baseline comparison important?
- What should remain constant for fair comparison?
Stretch (optional)
Save best parameters by validation loss.
Likely test focus
- Forward pass output dimensions.
- Training loop returns decreasing trend on toy data.
- Neural model beats bigram on provided validation set.
What should improve
You can now compare neural and count-based baselines under fixed evaluation controls.