Chapter 06 · Where it all converges

The transformer block, assembled

You now own every part. Dot products measure agreement (ch. 01). Matrices are learned functions (ch. 02). Softmax turns scores into shares (ch. 03). Cross-entropy trains it (ch. 04) via gradients (ch. 05). This chapter spends that savings account on the machine itself: attention, the famous equation read symbol by symbol, then the full block around it. The goal is a specific feeling: that the transformer is not a new kind of math, just the previous five chapters arranged with excellent taste.

Where this lives in a transformer · chapter 06
The idea
Attention is a soft, differentiable lookup: every token asks a question (query), advertises what it holds (key), and offers content (value); dot products score the matches and softmax shares out the reading.
Where you'll meet it
This is the transformer. Everything since 2017 (GPTs, Claude, Gemini, Llama) stacks this block, varying mostly the trims (chapter 07).
The one equation
Attention(Q,K,V) = softmax(QKᵀ/√d) · V
If you remember one thing
Each token's new vector is a weighted average of other tokens' value vectors, with weights = softmax of query·key agreements. Everything else in the block is support crew.

6.1The problem: one vector per token isn't enough

Chapter 01 gave every token a vector from a lookup table, and flagged the limitation: the table's vector for bank is identical in "river bank" and "bank deposit." Meaning depends on context, and a lookup table has none. What's needed is a mechanism that lets each token's vector be updated by the other tokens around it: selectively (the word river matters to bank; the word the mostly doesn't) and differentiably (chapter 05's button must reach through it). Selective, differentiable reading from the rest of the sentence: build that, and context flows. Attention is that mechanism, and its genius is to build "reading" out of nothing but the operations you already know.

6.2Queries, keys, values: a lookup made of dot products

Here is the equation that launched a thousand papers, annotated with everything the last five chapters taught you:

Plate 6·AThe attention equation, read with chapter numbers
softmax(QKT / √d) · V Q = X·W_Q (T×d) one learned matrix (ch. 02) turns each token's vector into a question: "what am I looking for right now?" KT: keys, transposed (d×T) a second learned matrix makes each token an advertisement: "here's what I hold." QKT = every question · every ad (ch. 01) ÷ √d keeps dot products softmax-sized (mathbox below) softmax, per row (ch. 03) each token's raw match scores become a budget over the tokens it may read from: shares ≥ 0, summing to 1 · V : spend the budget (ch. 02) a third learned matrix makes each token's content; multiplying by the weights takes a weighted AVERAGE of the values read
One more mechanism hides in decoder models: a causal mask sets every score where a token would peek at its own future to −∞ before the softmax (e−∞ = 0: a zero share, by chapter 03's math). Order enforced by arithmetic.

The dictionary analogy makes the names click. In a hash map, a query either matches a key exactly (fetch the value) or doesn't (fetch nothing). Attention is the soft version: every query matches every key to a degree (the dot product), and what comes back is a blend of all the values, weighted by match quality. Soft matters twice over: it lets partial relevance count, and it makes the whole lookup differentiable, so chapter 05's gradients flow through reading itself, and the model can learn what to look for. WQ, WK, WV are learned; nobody told the model that pronouns should query for their antecedents. Prediction pressure grew that, the same way it grew the embedding geometry.

6.3The grid: T² agreements, then a weighted average

Run the shape-check from chapter 02 and the structure falls out. Q is (T×d), Kᵀ is (d×T), so QKᵀ is (T×T): one score per pair of tokens. Row i holds token i's match scores against everyone it may read (with the causal mask: itself and earlier tokens). Softmax each row into a budget; multiply by V, and row i's output is a weighted average of the value vectors it chose, chapter 02's "recipe mixing rows" reading in the flesh.

Plate 6·BOne head reading "the robot picked up the ball because it was heavy"
attention weights after softmax · row = reader · column = being read · darker = bigger share robot picked up ball because it robot ▸ ball ▸ it ▸ .22 .04 .03 .58 .06 .07 ← 58% of "it"'s budget reads "ball" so the vector written back for "it" ≈ .58·v(ball) + .22·v(robot) + … : "it" now MEANS mostly-ball, and the next layer can act on that. Upper-right of the grid: masked to zero (no reading the future). Cost, honestly: T tokens → T² cells. Double the context, quadruple the grid. This bill drives ch. 07.
Illustrative weights, real phenomenon: pronoun-resolution heads like this are routinely found in trained models. Watch live ones form on your own sentences in the Transformer Explainer.
Why √d: the one scaling factor, derived in four sentences

A query and a key are d-dimensional (say d = 64), with entries that, at initialization, are roughly independent with mean 0 and variance 1. Their dot product is then a sum of d such products, so its variance is about d and its typical size is about √d ≈ 8.

Scores that size are typically also that far apart, and chapter 03's ratio rule prices the gaps: an 8-logit lead means one token outweighs another by e⁸ ≈ 3,000×. One token takes nearly the whole budget, and the distribution is saturated. Saturated softmax has near-zero gradients (nudging a logit barely moves the shares), so chapter 05's learning signal dies exactly where it's needed most.

Dividing by √d rescales typical scores back to size ~1, softmax stays in its responsive zone, gradients live. That is the whole story: the most famous denominator in machine learning is a variance correction protecting the learning signal.

6.4Many heads: parallel readers with different agendas

One budget per token is cramping: it may want to read for its antecedent and for the verb governing it, and a single softmax makes those desires compete. So the block runs several attentions in parallel: heads, each with its own small WQ, WK, WV (GPT-2 small: 12 heads of 64 dims each, slicing the 768). Each head maintains its own grid and its own learned agenda; trained models reliably grow heads tracking syntax, positions, rare tokens, matching brackets. Concatenate the heads' outputs, mix once more with a final learned matrix, and you have multi-head attention: a committee of cheap specialist readers in place of one expensive generalist: more distinct reading patterns for the same arithmetic budget, and (chapter 07 will lean on this) independent knobs for later architects to turn.

6.5The rest of the block: stream, norm, MLP

Zoom out one notch and the whole transformer appears: the same block, stacked (12 times in GPT-2 small, around a hundred in frontier models). Within each block, three support systems, each one of your chapters wearing work clothes:

Plate 6·COne block, complete: everything is a small edit added to a running sum
the residual stream · each token's running vector norm multi-head attention tokens read each other · §6.2–6.4 + norm MLP (768→3072→768) each token thinks alone · ch. 02 §2.5 + from block below to block above · ×12 in GPT-2 · ~×100 at the frontier note the +: outputs are ADDED onto the stream, never replacing it
Modern models put the norm before each branch (as drawn, "pre-norm"); the 2017 original normed after. One of many small trims that separate eras of the same design.

The residual stream (the red spine) is chapter 05's soapbox made structural. Because every branch adds its output to the stream instead of replacing it, the identity path gives gradients an untouched multiply-by-1 route from the loss to the earliest layers: the vanishing-gradient relay problem, dissolved by a plus sign. It also gives the whole model a clean mental architecture: a shared bus, per token, that each layer reads from and writes small edits onto; a hundred layers deep, meaning is the accumulated sum of a hundred nudges. Interpretability researchers take this bus picture literally (Anthropic's Mathematical Framework for Transformer Circuits built a subfield on it), and chapter 07 reads papers that treat it as the model's public API. LayerNorm is chapter 01's length/direction split turned into plumbing: before each branch, recenter each token's vector and rescale it to standard length (plus two small learned dials; RMSNorm, the modern default, skips the recentering and keeps only the rescale). Directions carry the meaning; the norm keeps a hundred layers of additions from blowing the magnitudes out of every downstream operator's comfortable range. The MLP you already met in §2.5: two matrices and a kink, the per-token "thinking" step, holder of two-thirds of each block's parameters. Attention moves information between tokens; the MLP processes it within each token. That division of labor is the whole block.

6.6Position: RoPE, or meaning that rotates

One last debt from chapter 01. Attention as described is order-blind: the dot products in QKᵀ care about what vectors contain, not where their tokens sit, and "dog bites man" must not equal "man bites dog." Modern models inject order with RoPE (rotary position embeddings), a trick elegant enough to enjoy purely as geometry: treat each query and key vector as a stack of 2-D pairs, and rotate each pair by an angle proportional to the token's position (different pairs spin at different frequencies, like clock hands from seconds to years). The payoff drops out of rotation algebra: rotate q by angle mθ and k by nθ, and their dot product's dependence on position collapses to m − n, the relative distance. "How far apart are we?" gets baked into every attention score, "position 7 versus position 12" stops mattering, and the rotation framing is what the field's context-stretching tricks (position interpolation and its successors) are built on, which is part of why RoPE (from Su et al., 2021) conquered the field. EleutherAI's Rotary Embeddings: A Relative Revolution tells the full story, derivation and code included.

6.7The best teachers for this

This chapter has the deepest bench of brilliant teaching on the internet. The shortlist:

video · 26 min Attention in transformers, step-by-step (DL Ch. 6) 3Blue1Brown Widely considered the finest visual explanation of attention in existence: Q/K/V, the grid, masking, multi-head, all animated on real GPT weights. visual essay The Illustrated Transformer Jay Alammar The classic step-by-step diagram walkthrough of this entire chapter; the mental pictures most practitioners carry came from this page. interactive · 3-D LLM Visualization Brendan Bycroft A working GPT rendered as an explorable 3-D structure: fly through every matrix of this chapter, drawn to scale, animated mid-inference. Unforgettable. article Rotary Embeddings: A Relative Revolution EleutherAI · Biderman et al. The definitive RoPE explainer, from the team that popularized it in open models: the relative-distance derivation of §6.6, plus working code.

What you should now believe

Go deeper

Contents · Glossary · The syllabus · Sources credited inline; links verified 2026-08-07.