Chapter 04 · The scoreboard

Loss, surprise, and information

Training a language model means minimizing one number. When a paper shows a loss curve bending downward, when a lab announces "loss 2.1 on the validation set," when someone quotes a perplexity, they are all reading the same scoreboard: cross-entropy. The formula looks arbitrary (−Σ p log q?) until you build it from one primitive idea, at which point it becomes the only formula you'd accept. The idea: how surprised were you?

Where this lives in a transformer · chapter 04
The idea
Surprise is −log p: cheap for outcomes you expected, brutal for ones you dismissed. A model's loss is its average surprise at the actual next token.
Where you'll meet it
The pretraining objective of every LLM ever shipped; every loss curve in every paper; "perplexity" in evaluations; the KL-divergence terms in RLHF and distillation papers.
The one equation
loss = −log q(the token that actually came next), averaged over the corpus. Perplexity = eloss.
If you remember one thing
Training an LLM is one instruction repeated forever: be less surprised by real text.

4.1Surprise, made a number

Suppose an event you assigned probability p actually happens. How surprised are you? Whatever number we invent should meet three common-sense demands: certainty means zero surprise (p = 1 → 0); rarer means more surprising (surprise grows as p shrinks); and independent surprises should add: getting two shocks of probability p and q (joint probability pq) should surprise you by the sum of the individual surprises. There is essentially one function with all three properties:

surprise(p) = −log p

The minus sign is bookkeeping: log of a probability (≤ 1) is negative, so we flip it to keep surprise positive. The additivity demand is what forces a logarithm: it's the function that turns products into sums.

Plate 4·AThe price of being wrong grows without bound
probability p you gave the thing that happened surprise −log₂ p (bits) p = 1: saw it coming, 0 bits p = ½: a coin flip, 1 bit p = 0.1: mildly stunned, ≈3.3 bits p = 0.03 and shrinking: the cliff as p → 0, surprise → ∞: there is no flat "maximum penalty," ever
Measured in bits because the log is base-2 here (papers use natural log; same shape, different unit, called nats). The unbounded left edge is why models are so strongly discouraged from writing off possibilities entirely: confident wrongness is priced like an unbounded penalty.

The unit has a famous pedigree. With base-2 logs, surprise is measured in bits, and this is precisely Shannon's information: an event of probability ½ carries 1 bit; probability 1/1024, 10 bits. "Information" and "surprise" are the same quantity read in two directions: the less expected the message, the more it tells you. This one identity is the whole reason "language modeling" and "compression" keep showing up in the same sentence, and Chris Olah's Visual Information Theory turns it into pictures better than anything else written.

4.2Entropy: a distribution's built-in suspense

Now point the surprise meter at a whole distribution. Entropy is the surprise you should expect, on average, when outcomes are drawn from a distribution p and scored by that same distribution:

H(p) = −Σᵢ pᵢ log pᵢ

Each outcome's surprise, weighted by how often it occurs.

Entropy measures how much suspense the distribution inherently carries. A certain outcome: 0 bits, no suspense. A fair coin: 1 bit. A fair 8-sided die: 3 bits. English text, per character, famously lands near 1 bit per charactersoapboxShannon's 1951 human-prediction experiments put printed English somewhere around 0.6–1.3 bits per character: astonishingly low against log₂(27) ≈ 4.75 for random letters, because English is mostly redundant. That gap between "raw symbols" and "actual suspense" is exactly the compressibility of language, and an LLM's whole job is to close in on it., far below random letters, because language is predictable. Entropy is the floor: no predictor of a source can average less surprise than the source's own entropy. Training pushes a model toward that floor; it can never tunnel below it. The residual suspense of language belongs to language, not to the model.

4.3Cross-entropy: the number on every loss curve

One substitution turns entropy into the training objective of every LLM. Let outcomes still be drawn from the true source p, but score the surprise using your model's beliefs q:

H(p, q) = −Σᵢ pᵢ log qᵢ   ≥   H(p)

Reality picks the outcomes; your model pays the surprise bill. The bill is smallest, and equals the entropy floor, exactly when q = p.

That inequality is the entire philosophy of training in one line: your average surprise is minimized by believing the truth. Any gap between your beliefs and reality shows up as extra surprise you pay on average, so "minimize cross-entropy" and "make q match p" are the same project. In LLM training, the recipe each step is concrete to the point of anticlimax: the training target for one position is just the token that actually came next (probability 1 on blue, 0 elsewhere), so the sum collapses and the loss for that position is simply −log q(blue): the model's surprise at the actual next token, the very number from Plate 4·A. Average over billions of positions; that average is the loss curve. (One bookkeeping care: a single position's target, being certain, has no suspense of its own. §4.2's floor lives in the average: paying −log q at billions of tokens drawn from real text is precisely how the bill comes to estimate the model's cross-entropy against language's true distribution, entropy floor and all.)

Soapbox: read a loss curve like an insider

Those celebrated plots of "loss vs. training compute" are average-surprise curves, and their units mean things: nats (or bits) of average surprise per token, and lower-loss generations of models are literally less surprised by text. The curves flatten because of §4.2: the entropy of language itself is the floor, and the closer you get, the more compute each remaining hundredth costs. When you hear "scaling laws," picture this: an asymptote priced in surprise.

4.4Perplexity: the same number, made humane

Cross-entropy in nats reads like a physicist's lab note. Exponentiate it and it becomes something you can feel: perplexity = eloss, the model's effective number of equally-likely choices per token. A perplexity of 20 means the model is, on average, as torn as if it were choosing uniformly among 20 plausible next tokens. Loss 3.0 ≈ perplexity 20; loss 2.3 ≈ perplexity 10. Same scoreboard, friendlier scale: this is why evaluation tables quote perplexity while training logs quote loss, and converting between them in your head (one exponential) is a tiny superpower when reading papers. You already met this quantity: the temperature widget in chapter 03 reported "effective choices" as you dragged the slider; that readout was 2entropy-in-bits, this section's math pointed at a single distribution.

4.5KL divergence: the gap itself, named

Subtract the floor from the bill and the leftover gets its own celebrated name, the KL divergence: DKL(p‖q) = H(p, q) − H(p): the extra surprise you pay for believing q when the truth is p. Zero exactly when the beliefs match, positive otherwise, and asymmetric (using q where p belongs is a different sin from the reverse). You now hold the full kit of three: entropy (the floor), cross-entropy (the bill), KL (the overcharge). File KL especially: when an RLHF paper says the tuned model is penalized for "drifting from the base policy," or a distillation paper matches a student to a teacher, the penalty term is a KL divergence, and you can now read it as: a leash, priced in extra bits of surprise.

The derivation: from "make the data likely" to cross-entropy

The most principled starting point is maximum likelihood: choose model weights that assign the training text the highest possible probability. By chapter 03's chain factorization, that probability is a product over positions: Πₜ q(xₜ | x₁…xₜ₋₁).

Products of thousands of tiny numbers are numerically hopeless and analytically clumsy, so take the log (monotonic, so the argmax doesn't move): Σₜ log q(xₜ | …). Maximizing that is minimizing its negation: −Σₜ log q(xₜ | …), which is exactly "total surprise at the actual tokens." Divide by the token count for the average, and you have arrived at the cross-entropy loss with no step along the way that involved taste. That is the sense in which the LLM objective is not one choice among many: it is what "make the data likely" turns into when you take a logarithm.

Perplexity bookkeeping, for completeness: PPL = eloss-in-nats = 2loss-in-bits. A uniform belief over N options has loss log N and perplexity exactly N, which is what licenses the reading "effective number of choices."

4.6The best teachers for this

video · 26 min The Key Equation Behind Probability Artem Kirsanov Builds surprise → entropy → cross-entropy → KL visually, from first principles, and lands exactly at "why this is the loss function of LLMs." The only video that derives the formula instead of asserting it. visual essay Visual Information Theory Christopher Olah Entropy as the length of an optimal code, cross-entropy as the cost of using the wrong code. A decade old and still the definitive picture of what "bits" buy you. article · 20 min Evaluation Metrics for Language Modeling Chip Huyen · The Gradient Connects this chapter directly to how models are scored in practice: entropy, cross-entropy, perplexity, bits-per-character, and what the numbers looked like for real models.

What you should now believe

Go deeper

Contents · Glossary · The syllabus · Sources credited inline; links verified 2026-08-07.