Built for reliable learning, not fast content

A hands-on curriculum with measurable quality controls

You learn by writing code, getting deterministic feedback, and improving step by step. Behind the scenes, lessons are checked against source research and every release must pass quality gates before it ships.

Curriculum scope is grounded in Attention Is All You Need (Vaswani et al., 2017). Faithfulness snapshots refresh daily and engineering quality snapshots refresh with successful mainline releases.

What We Built

This product teaches core language-model mechanics through short lessons and practical coding loops, not long lectures.

Hands-On Lessons

Every lesson asks you to implement one concrete behavior, run checks, and iterate until your solution works.

Deterministic Feedback

The same code produces the same result, so learners can trust the signal and focus on understanding instead of flaky tooling.

Progress That Sticks

Your code and attempt history are saved on your device by default, so you can continue where you left off without account friction.

Built for Reliability

The learning flow is backed by release guardrails so quality, performance, and accessibility are checked before updates land.

How Learning Works

Read, implement, test, and improve in short loops.

1. Learn one concept at a time with a short explanation and a concrete task.

2. Write code directly in the lesson and run checks with case-level feedback.

3. Fix issues quickly with deterministic reruns and clear pass/fail output.

4. Build confidence through repeated, predictable progress across the catalogue.

Why You Can Trust It

Reliability comes from evidence, reproducibility, and clear gates.

Claims are checked against source evidence from the paper, not ad-hoc summaries.

Learning is active: you prove understanding by shipping working code.

Checks are reproducible, which makes progress signals stable and explainable.

Automated and human review both contribute before content is trusted.

Releases are blocked if critical quality checks fail.

Curriculum Quality Approach

Lessons are designed for clarity first, then validated for consistency and coverage.

How Lessons Are Built and Reviewed

Content starts as plain-language lessons with concrete examples and practice tasks.

Automated validation checks structure, sequence, and exercise integrity so learners get a consistent experience.

We prioritize direct explanations and practical outcomes over broad but shallow coverage.

Engineering Quality Signals

Reliability is treated as a release requirement, not a best effort.

Quality Guardrails

Automated tests verify core learning and grading behavior before release.

Static checks enforce correctness, type safety, formatting, and maintainability limits.

Complexity guardrails catch code that becomes too hard to reason about.

Build verification ensures the production app can ship safely.

Performance, accessibility, and SEO audits run on key learner-facing pages.

If any required gate fails, release is blocked until it is fixed.

Snapshot Source

snapshot

Last Updated

2/9/2026, 2:45:51 PM

Lighthouse Assertion Failures

0

Line Coverage

93.8%

Coverage Snapshot

Coverage is a guarded quality signal, not a claim that every edge case in the product is fully tested.

Lines

93.8% (91/97)

Statements

93.8% (91/97)

Functions

100.0% (6/6)

Branches

79.3% (46/58)

Lighthouse Snapshot

Scores come from repeated representative runs on key learner pages.

RoutePerformanceAccessibilityBest PracticesSEO
/
100%
100%
100%
100%
/curriculum
100%
100%
100%
100%
/lesson/01-01
93%
100%
100%
100%

Live Faithfulness Dashboard

Current catalogue-level scores from the latest pipeline run.

Lessons Scored

18/32

Total Claims

63

Faithfulness Gate

PASS

Last Updated

2/9/2026, 3:56:30 AM

How Faithfulness Is Measured

Each lesson is split into factual claims and compared against evidence from the reference paper.

Measurement Method

1. Extract candidate claims from each lesson.

2. Retrieve the most relevant evidence passages from the paper.

3. Score each claim against that evidence.

4. Label claims as entailed, neutral, or contradicted using fixed thresholds.

5. Aggregate metrics and evaluate gate checks for coverage, consistency, and risk rates.

These metrics are strong signals, not proof of absolute truth, and can be affected by retrieval quality and threshold settings.

Metric Definitions

Core formulas used in the report.

Coverage = entailed / total claims

Consistency = (entailed - contradicted) / total claims

Hallucination rate = neutral / total claims

Contradiction rate = contradicted / total claims

Live Method Snapshot

Current scoring configuration for transparency.

Paper Source

aiayn.pdf

Target Files

32

Paper Chunks

27

Retrieval Top-K

5

Embedding Model

Xenova/all-MiniLM-L6-v2

NLI Model

Xenova/nli-deberta-v3-xsmall

Claim Label Thresholds

entailedMin: 0.015

contradictedMin: 0.999

contradictionMargin: 0.95

minSimilarityForSupport: 0.2

minSimilarityForContradiction: 0.25

Gate Thresholds

minCoverage: 0.2

minConsistency: 0.1

minEvidenceQuality: 0.35

maxHallucinationRate: 0.7

maxContradictionRate: 0.15

Coverage

85.7%

Consistency

0.857

Evidence Quality

0.356

Hallucination Rate

14.3%

Contradiction Rate

0.0%

Gate Checks

Threshold checks from the latest faithfulness run.

coverage: PASS
consistency: PASS
evidenceQuality: PASS
hallucinationRate: PASS
contradictionRate: PASS
highConfidenceUnsupportedRate: PASS

Score Legend

Status combines coverage, hallucination risk, contradiction risk, and data availability.

Good
Watch
Poor
No data

Unit Faithfulness

Unit-level quality includes claim depth, label mix, and status from the latest run.

Unit 01: First Wins (Coding Basics)

No data

0/4 lessons with claims · 0 claims

No claim data

No claims extracted in this run

Unit 02: First Working LM (Bigram)

No data

0/4 lessons with claims · 0 claims

No claim data

No claims extracted in this run

Unit 03: Trainable Neural LM

Good

1/4 lessons with claims · 2 claims

2 entailed
0 neutral
0 contradicted

100.0% coverage · 0.0% neutral

Unit 04: Embeddings and Better Context

Good

1/4 lessons with claims · 4 claims

4 entailed
0 neutral
0 contradicted

100.0% coverage · 0.0% neutral

Unit 05: Attention Fundamentals

Good

4/4 lessons with claims · 17 claims

14 entailed
3 neutral
0 contradicted

82.4% coverage · 17.6% neutral

Unit 06: Transformer Block

Good

4/4 lessons with claims · 5 claims

5 entailed
0 neutral
0 contradicted

100.0% coverage · 0.0% neutral

Unit 07: Tiny GPT Training Loop

Good

4/4 lessons with claims · 22 claims

20 entailed
2 neutral
0 contradicted

90.9% coverage · 9.1% neutral

Unit 08: Capstone Quality and Interpretability

Good

4/4 lessons with claims · 13 claims

9 entailed
4 neutral
0 contradicted

69.2% coverage · 30.8% neutral

Lesson Faithfulness

Lesson status shows confidence signals when claims are present, and explicit no-data states when no claims were extracted.

01-01

First Green Test

No data
No claims
No claim data

No claims extracted in this run

Open lesson

01-02

Counting What the Model Sees

No data
No claims
No claim data

No claims extracted in this run

Open lesson

01-03

Random Choice for Generation

No data
No claims
No claim data

No claims extracted in this run

Open lesson

01-04

Unigram Baseline Generator

No data
No claims
No claim data

No claims extracted in this run

Open lesson

02-01

Tokenize and Vocabulary

No data
No claims
No claim data

No claims extracted in this run

Open lesson

02-02

Build Bigram Counts

No data
No claims
No claim data

No claims extracted in this run

Open lesson

02-03

Bigram Probabilities and Predict

No data
No claims
No claim data

No claims extracted in this run

Open lesson

02-04

Generate with Bigram

No data
No claims
No claim data

No claims extracted in this run

Open lesson

03-01

Make Context Target Examples

No data
No claims
No claim data

No claims extracted in this run

Open lesson

03-02

One-Hot and Linear Logits

No data
No claims
No claim data

No claims extracted in this run

Open lesson

03-03

Softmax and Cross-Entropy

Good
2 claims
2 entailed
0 neutral
0 contradicted

100.0% coverage · 0.0% neutral

Open lesson

03-04

Gradient Step That Reduces Loss

No data
No claims
No claim data

No claims extracted in this run

Open lesson

04-01

Embedding Lookup

No data
No claims
No claim data

No claims extracted in this run

Open lesson

04-02

Combine Context Vectors

No data
No claims
No claim data

No claims extracted in this run

Open lesson

04-03

MLP Language Head

No data
No claims
No claim data

No claims extracted in this run

Open lesson

04-04

Train Neural Language Model

Good
4 claims
4 entailed
0 neutral
0 contradicted

100.0% coverage · 0.0% neutral

Open lesson

05-01

Query, Key, Value Intuition

Good
3 claims
3 entailed
0 neutral
0 contradicted

100.0% coverage · 0.0% neutral

Open lesson

05-02

Attention Scores and Weights

Good
5 claims
4 entailed
1 neutral
0 contradicted

80.0% coverage · 20.0% neutral

Open lesson

05-03

Weighted Sum of Values

Good
4 claims
4 entailed
0 neutral
0 contradicted

100.0% coverage · 0.0% neutral

Open lesson

05-04

Causal Mask (No Peeking)

Watch
5 claims
3 entailed
2 neutral
0 contradicted

60.0% coverage · 40.0% neutral

Open lesson

06-01

Multi-Head Attention

Good
1 claims
1 entailed
0 neutral
0 contradicted

100.0% coverage · 0.0% neutral

Open lesson

06-02

Feed-Forward Layer

Good
1 claims
1 entailed
0 neutral
0 contradicted

100.0% coverage · 0.0% neutral

Open lesson

06-03

Residual and LayerNorm

Good
1 claims
1 entailed
0 neutral
0 contradicted

100.0% coverage · 0.0% neutral

Open lesson

06-04

Full Transformer Block Forward Pass

Good
2 claims
2 entailed
0 neutral
0 contradicted

100.0% coverage · 0.0% neutral

Open lesson

07-01

Positional Information

Good
6 claims
5 entailed
1 neutral
0 contradicted

83.3% coverage · 16.7% neutral

Open lesson

07-02

Stack Blocks Into MiniGPT

Good
5 claims
4 entailed
1 neutral
0 contradicted

80.0% coverage · 20.0% neutral

Open lesson

07-03

Batching and Train Step

Good
7 claims
7 entailed
0 neutral
0 contradicted

100.0% coverage · 0.0% neutral

Open lesson

07-04

Validation and Checkpoints

Good
4 claims
4 entailed
0 neutral
0 contradicted

100.0% coverage · 0.0% neutral

Open lesson

08-01

Temperature Sampling

Good
4 claims
3 entailed
1 neutral
0 contradicted

75.0% coverage · 25.0% neutral

Open lesson

08-02

Top-k Sampling

Good
2 claims
2 entailed
0 neutral
0 contradicted

100.0% coverage · 0.0% neutral

Open lesson

08-03

Visualize Attention Maps

Poor
2 claims
0 entailed
2 neutral
0 contradicted

0.0% coverage · 100.0% neutral

Open lesson

08-04

Final Model Comparison

Good
5 claims
4 entailed
1 neutral
0 contradicted

80.0% coverage · 20.0% neutral

Open lesson

Known Gaps and What's Next

We publish limitations directly so learners can judge progress honestly.

Faithfulness scoring continues to improve, and some valid teaching claims can still be marked unsupported when retrieval misses context.

Coverage metrics represent a guarded test scope, not exhaustive verification of every edge case in the product.

Some advanced topics are intentionally simplified for beginners; precision notes are added where those simplifications matter.

Faithfulness snapshots refresh daily and quality snapshots refresh with successful mainline releases.