Chapter 03 · The output side
An LLM never outputs a word. It outputs a probability distribution: a belief, spread across its entire vocabulary, about what token comes next. Every "the model said" you have ever read was really "we picked from the model's distribution," usually by rolling dice against it. This chapter rebuilds the two pieces of math that make that sentence precise: softmax, which turns raw scores into honest probabilities, and the sampling dials (temperature, top-k, top-p) that every API exposes and most engineers set by folklore.
Everything this chapter needs from probability theory fits in one paragraph. A (discrete)
probability distribution over N outcomes is a list of N numbers that are each ≥ 0 and that sum to
exactly 1. That's it: a budget of belief, one share per outcome. "The model
thinks blue is likely" means the model's budget hands blue a big share.
Certainty is the entire budget on one outcome; total ignorance is the budget split evenly
(1/N each, the uniform distribution). Between those extremes lies every distribution an
LLM will ever produce, and chapter 04 will measure exactly where between them a belief
sits. If you want to feel these definitions in your hands before moving on, Brown University's
Seeing Theory is a beautiful half hour.
Rewind to what the network actually computes. After the last block, the model holds one final vector (768 numbers in GPT-2) summarizing "everything relevant to what comes next." The output layer then plays chapter 01's game at full scale: dot-product that vector against every token's direction in the vocabulary, producing 50,257 agreement scores. Those raw scores are the logits.
Logits are useful but not yet a belief. They can be negative; they sum to whatever they happen
to sum to; a logit of 5.1 means nothing by itself. Their meaning is entirely
relative: blue at 7.2 versus green at 5.2 says blue has
two units more evidence, and (after this chapter's math) you can even say precisely how much more
probable that makes it: e² ≈ 7.4 times. What's needed is a converter that respects the rankings
and the gaps while producing an honest budget: non-negative shares, summing to 1.
The converter the entire field settled on is two moves long:
Exponentiate every logit (all results now positive), then divide each by the grand total (now they sum to 1).
Move one: ez. The exponential maps any number, however negative, to something positive, and it does so monotonically: higher logit, higher result, always. Rankings survive. Move two: divide by the sum. Anything divided by its own total sums to 1. Both requirements of a budget, met by construction, with the gaps between logits turned into ratios between probabilities.
Why the exponential specifically, and not something tamer like "just shift everything positive and divide"? Three defensible answers, in increasing depth. Practical: exp makes big logits dominate decisively, so the function behaves like a soft, differentiable version of "pick the max" (hence the name softmax; chapter 05 explains why differentiable matters). Structural: exp is the function (unique up to its base, which is exactly the dial temperature turns) that converts logit addition into probability multiplication, so evidence adds while beliefs multiply, which is how probabilistic evidence is supposed to compose. And a pedigree note: this is the Boltzmann distributionsoapboxIn statistical physics, the probability a system sits in a state of energy E is proportional to e−E/T, with T the temperature. Softmax-with-temperature is this equation wearing a machine-learning costume, and that is literally where the parameter's name comes from. from physics, where the "temperature" you're about to meet got its name.
Every LLM API exposes a knob called temperature, and its entire implementation is one division. Before the exponential, scale the logits: z → z/T. That's all it is.
Shift invariance. Add any constant c to every logit and softmax doesn't change: e(zᵢ+c)/Σe(zⱼ+c) = ecezᵢ/(ecΣezⱼ), and the ec cancels. Only differences between logits matter. Every real implementation exploits this by subtracting max(z) first, so the biggest exponent is e⁰ = 1 and nothing overflows: when a paper or codebase says "numerically stable softmax," this is the whole trick.
What temperature does to ratios. The ratio between two probabilities is pᵢ/pⱼ = e(zᵢ−zⱼ)/T. At T=1, a 2-logit gap means e² ≈ 7.4×. At T=0.5 the same gap means e⁴ ≈ 55×; at T=2, e¹ ≈ 2.7×. Temperature reprices every gap by the same rule, which is why it feels like one coherent "confidence" dial rather than 50,257 separate ones.
Given the budget, something still has to pick. The menu you'll actually meet in an API, in one place:
| Strategy | Rule | Character |
|---|---|---|
| Greedy | always take the biggest share | Deterministic; safe; prone to loops and dull phrasing. What T→0 converges to. |
| Temperature | rescale logits by 1/T, then roll the dice | The one continuous dial: below 1 sharpens toward greedy, above 1 flattens toward chaos. |
| Top-k | keep only the k biggest shares, renormalize, roll | A hard guardrail: the long tail of barely-plausible tokens simply can't be drawn. |
| Top-p (nucleus) | keep the smallest set of tokens whose shares total p (say 0.9), renormalize, roll | An adaptive guardrail: the kept set shrinks when the model is confident, widens when it's torn. Usually composed with temperature. |
Two intuitions worth keeping. First, the guardrails exist because of softmax's one honest
weakness: everything gets a nonzero share, even tuba after "the sky is,"
and across thousands of generated tokens those slivers add up to occasional nonsense; truncation
is the fix. Second, there is no "correct" setting: sampling is a taste dial trading reliability
against variety, which is why code assistants ship near T=0 and brainstorming tools don't.
What you should now believe
Go deeper
Contents · Glossary · The syllabus · Sources credited inline; links verified 2026-08-07.