Lesson 08-04

Final Model Comparison

15 min
1 export
3 tests

Lesson blocked by prerequisites

Complete and save a passing attempt for your active lesson before running this one.

Go to active lesson

Lesson workspace sections

Submission + Results
Not run yet

Seed 804 • Runtime includes 31 prerequisite modules
Test Results

Run tests to see case-by-case feedback.

Attempts0 saved

No saved attempts yet.

Lesson README

08-04 Final Model Comparison

Why this matters

A fair capstone comparison keeps prompt, seed, and evaluation protocol fixed across model variants.

Intuition first (no jargon)

Compare architecture changes by pairing metric deltas with concrete output samples.

Paper grounding

  • The paper compares model variants with controlled ablations (for example attention heads and dimensional settings).
  • Fair comparison requires fixed evaluation protocol so changes in metrics can be attributed to architecture.

Worked example

  • Shapes:
    • report entries list [numModels]
    • each entry includes { modelName, metrics, sample, reflection }
  • One computed step:
    • if bigram validation loss is 2.45 and MiniGPT is 1.92, improvement is 0.53
  • In the paper, Table 3 studies head count and key/value dimensions while keeping computation roughly fixed.

Code walkthrough

js
export function compareModels(models, prompt, cfg) {
  // returns structured report
}

Your task

Implement compareModels(models, prompt, cfg).

  • Generate text from unigram, bigram, neural LM, and MiniGPT.
  • Compute comparable metrics (loss or surprise).
  • Return a report object ready for UI display.
  • For each model, include reflection in this format:
    • Change: what architectural capability was added.
    • Evidence: metric delta and one output example.
    • Limitation: what still fails.

Common mistakes

  • Different prompts or seeds across models

Precision note

Curriculum comparisons on tiny datasets can be noisy and small-sample biased. Directional evidence under fixed controls.

Hints

  • Use fixed seed and same prompt for fairness.
  • Keep generation length consistent.
  • Include both numeric metrics and short sample strings.

Check your thinking

  1. Which improvements came from attention specifically?
  2. Why is fairness setup critical in comparison?
  3. What remains limited in a tiny model?

Stretch (optional)

Add an "ablation" mode that disables one component and reruns comparison.

Likely test focus

  • Report schema correctness.
  • Presence of all model variants.
  • Reproducible comparison under fixed seed.

What should improve

You can now produce an evidence-backed final comparison report across model versions.

Bridge to next lesson

Curriculum complete: baseline counting through transformer-style generation.

Monaco Editor

Matches starter

Files

Editor is deferred on smaller screens to keep startup fast.

Autosave is enabled in local storage for this lesson.