Chapter 06 · Where it all converges
You now own every part. Dot products measure agreement (ch. 01). Matrices are learned functions (ch. 02). Softmax turns scores into shares (ch. 03). Cross-entropy trains it (ch. 04) via gradients (ch. 05). This chapter spends that savings account on the machine itself: attention, the famous equation read symbol by symbol, then the full block around it. The goal is a specific feeling: that the transformer is not a new kind of math, just the previous five chapters arranged with excellent taste.
Chapter 01 gave every token a vector from a lookup table, and flagged the limitation: the
table's vector for bank is identical in "river bank" and "bank deposit." Meaning
depends on context, and a lookup table has none. What's needed is a mechanism that lets each
token's vector be updated by the other tokens around it: selectively (the word
river matters to bank; the word the mostly doesn't) and
differentiably (chapter 05's button must reach through it). Selective, differentiable reading
from the rest of the sentence: build that, and context flows. Attention is that mechanism, and
its genius is to build "reading" out of nothing but the operations you already know.
Here is the equation that launched a thousand papers, annotated with everything the last five chapters taught you:
The dictionary analogy makes the names click. In a hash map, a query either matches a key exactly (fetch the value) or doesn't (fetch nothing). Attention is the soft version: every query matches every key to a degree (the dot product), and what comes back is a blend of all the values, weighted by match quality. Soft matters twice over: it lets partial relevance count, and it makes the whole lookup differentiable, so chapter 05's gradients flow through reading itself, and the model can learn what to look for. WQ, WK, WV are learned; nobody told the model that pronouns should query for their antecedents. Prediction pressure grew that, the same way it grew the embedding geometry.
Run the shape-check from chapter 02 and the structure falls out. Q is (T×d), Kᵀ is (d×T), so QKᵀ is (T×T): one score per pair of tokens. Row i holds token i's match scores against everyone it may read (with the causal mask: itself and earlier tokens). Softmax each row into a budget; multiply by V, and row i's output is a weighted average of the value vectors it chose, chapter 02's "recipe mixing rows" reading in the flesh.
A query and a key are d-dimensional (say d = 64), with entries that, at initialization, are roughly independent with mean 0 and variance 1. Their dot product is then a sum of d such products, so its variance is about d and its typical size is about √d ≈ 8.
Scores that size are typically also that far apart, and chapter 03's ratio rule prices the gaps: an 8-logit lead means one token outweighs another by e⁸ ≈ 3,000×. One token takes nearly the whole budget, and the distribution is saturated. Saturated softmax has near-zero gradients (nudging a logit barely moves the shares), so chapter 05's learning signal dies exactly where it's needed most.
Dividing by √d rescales typical scores back to size ~1, softmax stays in its responsive zone, gradients live. That is the whole story: the most famous denominator in machine learning is a variance correction protecting the learning signal.
One budget per token is cramping: it may want to read for its antecedent
and for the verb governing it, and a single softmax makes those desires compete. So the
block runs several attentions in parallel:
heads, each with its own small WQ,
WK, WV (GPT-2 small: 12 heads of 64 dims each, slicing the 768). Each head
maintains its own grid and its own learned agenda; trained models reliably grow heads tracking
syntax, positions, rare tokens, matching brackets. Concatenate the heads' outputs, mix once more
with a final learned matrix, and you have multi-head attention: a committee of cheap
specialist readers in place of one expensive generalist: more distinct reading patterns
for the same arithmetic budget, and
(chapter 07 will lean on this) independent knobs for later architects to turn.
Zoom out one notch and the whole transformer appears: the same block, stacked (12 times in GPT-2 small, around a hundred in frontier models). Within each block, three support systems, each one of your chapters wearing work clothes:
The residual stream (the red spine) is chapter 05's soapbox made structural. Because every branch adds its output to the stream instead of replacing it, the identity path gives gradients an untouched multiply-by-1 route from the loss to the earliest layers: the vanishing-gradient relay problem, dissolved by a plus sign. It also gives the whole model a clean mental architecture: a shared bus, per token, that each layer reads from and writes small edits onto; a hundred layers deep, meaning is the accumulated sum of a hundred nudges. Interpretability researchers take this bus picture literally (Anthropic's Mathematical Framework for Transformer Circuits built a subfield on it), and chapter 07 reads papers that treat it as the model's public API. LayerNorm is chapter 01's length/direction split turned into plumbing: before each branch, recenter each token's vector and rescale it to standard length (plus two small learned dials; RMSNorm, the modern default, skips the recentering and keeps only the rescale). Directions carry the meaning; the norm keeps a hundred layers of additions from blowing the magnitudes out of every downstream operator's comfortable range. The MLP you already met in §2.5: two matrices and a kink, the per-token "thinking" step, holder of two-thirds of each block's parameters. Attention moves information between tokens; the MLP processes it within each token. That division of labor is the whole block.
One last debt from chapter 01. Attention as described is order-blind: the dot products in QKᵀ care about what vectors contain, not where their tokens sit, and "dog bites man" must not equal "man bites dog." Modern models inject order with RoPE (rotary position embeddings), a trick elegant enough to enjoy purely as geometry: treat each query and key vector as a stack of 2-D pairs, and rotate each pair by an angle proportional to the token's position (different pairs spin at different frequencies, like clock hands from seconds to years). The payoff drops out of rotation algebra: rotate q by angle mθ and k by nθ, and their dot product's dependence on position collapses to m − n, the relative distance. "How far apart are we?" gets baked into every attention score, "position 7 versus position 12" stops mattering, and the rotation framing is what the field's context-stretching tricks (position interpolation and its successors) are built on, which is part of why RoPE (from Su et al., 2021) conquered the field. EleutherAI's Rotary Embeddings: A Relative Revolution tells the full story, derivation and code included.
This chapter has the deepest bench of brilliant teaching on the internet. The shortlist:
What you should now believe
Go deeper
Contents · Glossary · The syllabus · Sources credited inline; links verified 2026-08-07.