Chapter 03 · The output side

From scores to probabilities

An LLM never outputs a word. It outputs a probability distribution: a belief, spread across its entire vocabulary, about what token comes next. Every "the model said" you have ever read was really "we picked from the model's distribution," usually by rolling dice against it. This chapter rebuilds the two pieces of math that make that sentence precise: softmax, which turns raw scores into honest probabilities, and the sampling dials (temperature, top-k, top-p) that every API exposes and most engineers set by folklore.

Where this lives in a transformer · chapter 03
The idea
Beliefs are budgets: a distribution assigns every possible outcome a non-negative share, and the shares sum to exactly 1. Softmax is the standard machine for converting arbitrary scores into such a budget.
Where you'll meet it
In two roles, and the same math in both: at the very top, turning the final vector's 50,257 vocabulary scores into next-token probabilities; and inside every attention head (chapter 06), turning token-to-token match scores into mixing weights.
The one equation
softmax(z)ᵢ = ezᵢ / Σⱼ ezⱼ : exponentiate every score, then divide each by the total.
If you remember one thing
Logits are relative evidence; softmax is exponentiate-and-normalize; temperature just rescales the logits before the exponential, sharpening or flattening the same underlying preferences.

3.1A thirty-second probability refresher

Everything this chapter needs from probability theory fits in one paragraph. A (discrete) probability distribution over N outcomes is a list of N numbers that are each ≥ 0 and that sum to exactly 1. That's it: a budget of belief, one share per outcome. "The model thinks blue is likely" means the model's budget hands blue a big share. Certainty is the entire budget on one outcome; total ignorance is the budget split evenly (1/N each, the uniform distribution). Between those extremes lies every distribution an LLM will ever produce, and chapter 04 will measure exactly where between them a belief sits. If you want to feel these definitions in your hands before moving on, Brown University's Seeing Theory is a beautiful half hour.

3.2Logits: the scores before the honesty

Rewind to what the network actually computes. After the last block, the model holds one final vector (768 numbers in GPT-2) summarizing "everything relevant to what comes next." The output layer then plays chapter 01's game at full scale: dot-product that vector against every token's direction in the vocabulary, producing 50,257 agreement scores. Those raw scores are the logits.

Logits are useful but not yet a belief. They can be negative; they sum to whatever they happen to sum to; a logit of 5.1 means nothing by itself. Their meaning is entirely relative: blue at 7.2 versus green at 5.2 says blue has two units more evidence, and (after this chapter's math) you can even say precisely how much more probable that makes it: e² ≈ 7.4 times. What's needed is a converter that respects the rankings and the gaps while producing an honest budget: non-negative shares, summing to 1.

3.3Softmax: exponentiate, then share out

The converter the entire field settled on is two moves long:

softmax(z)ᵢ = ezᵢΣⱼ ezⱼ

Exponentiate every logit (all results now positive), then divide each by the grand total (now they sum to 1).

Move one: ez. The exponential maps any number, however negative, to something positive, and it does so monotonically: higher logit, higher result, always. Rankings survive. Move two: divide by the sum. Anything divided by its own total sums to 1. Both requirements of a budget, met by construction, with the gaps between logits turned into ratios between probabilities.

Plate 3·AThe two-move pipeline: scores → all-positive → budget
logits z 3.2 2.0 −1.5 1.0 blueredtubagreen any sign · any total eᶻ exponentiated 24.5 7.4 0.2 2.7 all positive · ranking intact · gaps became ratios ÷ 34.8 (the sum) probabilities .70 .21 .01 .08 non-negative · sums to 1 · a belief
Four-token toy vocabulary; a real model runs this same pipeline over 50,000+ logits, every single token it generates. Note what survived the pipeline: the order and the gaps. Note what appeared: honesty (a true budget).

Why the exponential specifically, and not something tamer like "just shift everything positive and divide"? Three defensible answers, in increasing depth. Practical: exp makes big logits dominate decisively, so the function behaves like a soft, differentiable version of "pick the max" (hence the name softmax; chapter 05 explains why differentiable matters). Structural: exp is the function (unique up to its base, which is exactly the dial temperature turns) that converts logit addition into probability multiplication, so evidence adds while beliefs multiply, which is how probabilistic evidence is supposed to compose. And a pedigree note: this is the Boltzmann distributionsoapboxIn statistical physics, the probability a system sits in a state of energy E is proportional to e−E/T, with T the temperature. Softmax-with-temperature is this equation wearing a machine-learning costume, and that is literally where the parameter's name comes from. from physics, where the "temperature" you're about to meet got its name.

3.4Temperature: one dial, same beliefs

Every LLM API exposes a knob called temperature, and its entire implementation is one division. Before the exponential, scale the logits: z → z/T. That's all it is.

The temperature dialinteractive · drag the slider

Eight fixed logits (a toy next-token belief after "The sky is"); only T moves. Watch how the same scores yield near-certainty or near-chaos. For the real thing, drag the identical slider on a live GPT-2 inside the Transformer Explainer.

Two lines of algebra every implementation relies on

Shift invariance. Add any constant c to every logit and softmax doesn't change: e(zᵢ+c)/Σe(zⱼ+c) = ecezᵢ/(ecΣezⱼ), and the ec cancels. Only differences between logits matter. Every real implementation exploits this by subtracting max(z) first, so the biggest exponent is e⁰ = 1 and nothing overflows: when a paper or codebase says "numerically stable softmax," this is the whole trick.

What temperature does to ratios. The ratio between two probabilities is pᵢ/pⱼ = e(zᵢ−zⱼ)/T. At T=1, a 2-logit gap means e² ≈ 7.4×. At T=0.5 the same gap means e⁴ ≈ 55×; at T=2, e¹ ≈ 2.7×. Temperature reprices every gap by the same rule, which is why it feels like one coherent "confidence" dial rather than 50,257 separate ones.

3.5Sampling: how a belief becomes a word

Given the budget, something still has to pick. The menu you'll actually meet in an API, in one place:

StrategyRuleCharacter
Greedyalways take the biggest shareDeterministic; safe; prone to loops and dull phrasing. What T→0 converges to.
Temperaturerescale logits by 1/T, then roll the diceThe one continuous dial: below 1 sharpens toward greedy, above 1 flattens toward chaos.
Top-kkeep only the k biggest shares, renormalize, rollA hard guardrail: the long tail of barely-plausible tokens simply can't be drawn.
Top-p (nucleus)keep the smallest set of tokens whose shares total p (say 0.9), renormalize, rollAn adaptive guardrail: the kept set shrinks when the model is confident, widens when it's torn. Usually composed with temperature.

Two intuitions worth keeping. First, the guardrails exist because of softmax's one honest weakness: everything gets a nonzero share, even tuba after "the sky is," and across thousands of generated tokens those slivers add up to occasional nonsense; truncation is the fix. Second, there is no "correct" setting: sampling is a taste dial trading reliability against variety, which is why code assistants ship near T=0 and brainstorming tools don't.

3.6The best teachers for this

interactive · live GPT-2 Transformer Explainer Polo Club · Georgia Tech A real GPT-2 running in your browser: type a prompt, open the output panel, and drag temperature and top-k/top-p while the actual distribution reshapes. This chapter, made tactile. video · 14 min Neural Networks Part 5: ArgMax and SoftMax StatQuest · Josh Starmer The slowest, most concrete walk through softmax arithmetic on video. The right "let me double-check I really got it" companion. article · short How do temperature, top-k, and top-p differ? Sebastian Raschka Exactly §3.4–§3.5 with a worked five-token example, from one of the field's most trusted explainers. The reference to send colleagues.

What you should now believe

Go deeper

Contents · Glossary · The syllabus · Sources credited inline; links verified 2026-08-07.