08-01 Temperature Sampling
Why this matters
Temperature rescales logits during sampling to control distribution sharpness without retraining weights.
Intuition first (no jargon)
Lower temperature sharpens probabilities; higher temperature flattens them.
Paper grounding
- The paper reports beam-search decoding (beam size 4) with length normalization during translation inference.
- This lesson adds temperature sampling as a curriculum extension for controllable generation behavior.
Worked example
- Inputs: logits
[2, 1, 0] - Shapes: logits
[V] -> scaled logits [V] -> probs [V] - One computed step:
T = 0.5-> sharper distributionT = 2.0-> flatter distribution
Code walkthrough
jsexport function sampleWithTemperature(logits, temperature, rng = Math.random) {}
sampleWithTemperature: rescale logits -> softmax -> sample token ID
Your task
Implement temperature-based sampling.
- Scale logits by dividing by temperature before softmax.
- If temperature is near zero, fallback to argmax.
- Sample token ID from resulting distribution.
Common mistakes
- Wrong scaling direction (
*Tinstead of/T) - No clamp near zero temperature
Precision note
Paper decoding reports beam search; this lesson explores temperature-based sampling control.
Hints
- Clamp minimum temperature to a tiny positive value.
- Reuse stable softmax and weighted sampling helpers.
- Compare output diversity at
0.7,1.0, and1.5.
Check your thinking
- Why does lower temperature reduce randomness?
- Why can high temperature produce gibberish?
- Why include argmax fallback?
Stretch (optional)
Expose temperature schedule during generation.
Likely test focus
- Correct argmax behavior near zero temperature.
- Valid sampled IDs.
- Distribution changes as temperature changes.
What should improve
You can now tune output diversity by adjusting only temperature.
Bridge to next lesson
Next lesson: combine temperature with candidate filtering.