Chapter 04 · The scoreboard
Training a language model means minimizing one number. When a paper shows a loss curve bending downward, when a lab announces "loss 2.1 on the validation set," when someone quotes a perplexity, they are all reading the same scoreboard: cross-entropy. The formula looks arbitrary (−Σ p log q?) until you build it from one primitive idea, at which point it becomes the only formula you'd accept. The idea: how surprised were you?
Suppose an event you assigned probability p actually happens. How surprised are you? Whatever number we invent should meet three common-sense demands: certainty means zero surprise (p = 1 → 0); rarer means more surprising (surprise grows as p shrinks); and independent surprises should add: getting two shocks of probability p and q (joint probability pq) should surprise you by the sum of the individual surprises. There is essentially one function with all three properties:
The minus sign is bookkeeping: log of a probability (≤ 1) is negative, so we flip it to keep surprise positive. The additivity demand is what forces a logarithm: it's the function that turns products into sums.
The unit has a famous pedigree. With base-2 logs, surprise is measured in bits, and this is precisely Shannon's information: an event of probability ½ carries 1 bit; probability 1/1024, 10 bits. "Information" and "surprise" are the same quantity read in two directions: the less expected the message, the more it tells you. This one identity is the whole reason "language modeling" and "compression" keep showing up in the same sentence, and Chris Olah's Visual Information Theory turns it into pictures better than anything else written.
Now point the surprise meter at a whole distribution. Entropy is the surprise you should expect, on average, when outcomes are drawn from a distribution p and scored by that same distribution:
Each outcome's surprise, weighted by how often it occurs.
Entropy measures how much suspense the distribution inherently carries. A certain outcome: 0 bits, no suspense. A fair coin: 1 bit. A fair 8-sided die: 3 bits. English text, per character, famously lands near 1 bit per charactersoapboxShannon's 1951 human-prediction experiments put printed English somewhere around 0.6–1.3 bits per character: astonishingly low against log₂(27) ≈ 4.75 for random letters, because English is mostly redundant. That gap between "raw symbols" and "actual suspense" is exactly the compressibility of language, and an LLM's whole job is to close in on it., far below random letters, because language is predictable. Entropy is the floor: no predictor of a source can average less surprise than the source's own entropy. Training pushes a model toward that floor; it can never tunnel below it. The residual suspense of language belongs to language, not to the model.
One substitution turns entropy into the training objective of every LLM. Let outcomes still be drawn from the true source p, but score the surprise using your model's beliefs q:
Reality picks the outcomes; your model pays the surprise bill. The bill is smallest, and equals the entropy floor, exactly when q = p.
That inequality is the entire philosophy of training in one line: your average surprise
is minimized by believing the truth. Any gap between your beliefs and reality shows up as
extra surprise you pay on average, so "minimize cross-entropy" and "make q match p" are the same
project. In LLM training, the recipe each step is concrete to the point of anticlimax: the
training target for one position is just the token that actually came next (probability 1
on blue, 0 elsewhere), so the sum collapses and the loss for that position is simply
−log q(blue): the model's surprise at the actual next token, the
very number from Plate 4·A. Average over billions of positions; that average is the loss curve.
(One bookkeeping care: a single position's target, being certain, has no suspense of its own.
§4.2's floor lives in the average: paying −log q at billions of tokens drawn from real
text is precisely how the bill comes to estimate the model's cross-entropy against language's
true distribution, entropy floor and all.)
Those celebrated plots of "loss vs. training compute" are average-surprise curves, and their units mean things: nats (or bits) of average surprise per token, and lower-loss generations of models are literally less surprised by text. The curves flatten because of §4.2: the entropy of language itself is the floor, and the closer you get, the more compute each remaining hundredth costs. When you hear "scaling laws," picture this: an asymptote priced in surprise.
Cross-entropy in nats reads like a physicist's lab note. Exponentiate it and it becomes something you can feel: perplexity = eloss, the model's effective number of equally-likely choices per token. A perplexity of 20 means the model is, on average, as torn as if it were choosing uniformly among 20 plausible next tokens. Loss 3.0 ≈ perplexity 20; loss 2.3 ≈ perplexity 10. Same scoreboard, friendlier scale: this is why evaluation tables quote perplexity while training logs quote loss, and converting between them in your head (one exponential) is a tiny superpower when reading papers. You already met this quantity: the temperature widget in chapter 03 reported "effective choices" as you dragged the slider; that readout was 2entropy-in-bits, this section's math pointed at a single distribution.
Subtract the floor from the bill and the leftover gets its own celebrated name, the KL divergence: DKL(p‖q) = H(p, q) − H(p): the extra surprise you pay for believing q when the truth is p. Zero exactly when the beliefs match, positive otherwise, and asymmetric (using q where p belongs is a different sin from the reverse). You now hold the full kit of three: entropy (the floor), cross-entropy (the bill), KL (the overcharge). File KL especially: when an RLHF paper says the tuned model is penalized for "drifting from the base policy," or a distillation paper matches a student to a teacher, the penalty term is a KL divergence, and you can now read it as: a leash, priced in extra bits of surprise.
The most principled starting point is maximum likelihood: choose model weights that assign the training text the highest possible probability. By chapter 03's chain factorization, that probability is a product over positions: Πₜ q(xₜ | x₁…xₜ₋₁).
Products of thousands of tiny numbers are numerically hopeless and analytically clumsy, so take the log (monotonic, so the argmax doesn't move): Σₜ log q(xₜ | …). Maximizing that is minimizing its negation: −Σₜ log q(xₜ | …), which is exactly "total surprise at the actual tokens." Divide by the token count for the average, and you have arrived at the cross-entropy loss with no step along the way that involved taste. That is the sense in which the LLM objective is not one choice among many: it is what "make the data likely" turns into when you take a logarithm.
Perplexity bookkeeping, for completeness: PPL = eloss-in-nats = 2loss-in-bits. A uniform belief over N options has loss log N and perplexity exactly N, which is what licenses the reading "effective number of choices."
What you should now believe
Go deeper
Contents · Glossary · The syllabus · Sources credited inline; links verified 2026-08-07.