Chapter 02 · Retrieval from earlier tokens
Attention lets a token use information from earlier tokens. It compares the current token with each one, then combines their information according to how well they match. Because each earlier token has its own slot, the model can retrieve a specific part of the context. Chapters 04–07 examine ways to reduce the cost of that lookup.
Every token's vector is projected three ways, and the three projections play different roles. The same colors identify them throughout the explainer:
First, score every query against every key with a dot productcontext Dot products are efficient on GPUs: each dimension takes a multiply-add, and many scores can be computed as a matrix multiplication.. Next, softmaxcontext Softmax exponentiates the scores and divides by their sum. It increases the influence of better matches while keeping all weights positive and summing to 1. The slider below shows how changing score sharpness affects those weights. turns the scores into weights. Finally, use those weights to combine the values:
M is the causal mask (§2.2); dh is the per-head dimension (64 in GPT-2).
In plain words, one attention step is five moves: ask (form a query), advertise (every past token exposes a key), score (dot products measure match), normalize (softmax turns scores into a probability-like distribution), blend (weighted sum of values).
Three tokens, the cat sat, and the query comes from sat. Suppose the
scaled dot products of sat's query against the three
keys came out to 2, 1, 0 (three plain numbers; how they
arose is the projections' business). Softmax exponentiates and normalizes:
scores 2 1 0
e^score 7.389 2.718 1.000 (sum = 11.107)
weights 0.665 0.245 0.090 (each e^score / sum, they sum to 1)
Give the three tokens simple 2-dim values, v(the) = (10, 0), v(cat) = (0, 10), v(sat) = (4, 4), and blend:
output = 0.665·(10, 0) + 0.245·(0, 10) + 0.090·(4, 4)
= (6.65, 0) + (0, 2.45) + (0.36, 0.36)
= (7.01, 2.81)
The output at sat is (7.01, 2.81). The value from the has the
largest weight, followed by cat and sat. A one-point difference
between two raw scores produces about a 2.7× difference in weight. Real attention heads use
learned projections and more dimensions, but the calculation is the same.
Each token needs three projections because looking for information, being found, and providing information are different jobs:
it,
as a seeker, looks for the noun it refers to; animal needs to be findable as a noun.
What it looks for is not the
same as what it offers, and a single shared vector would make the raw compatibility
score symmetric ("A matches B" ≡ "B matches A") before masking and normalization, when
language routinely wants it lopsided (an adjective seeks its noun far more than the noun seeks the
adjective). Separate Q and K let the model learn that lopsidedness directly.So the three projection matrices WQ, WK, WV are three learned questions asked of the same vector: what do I seek, how am I found, what do I give. Separate queries and keys make the lookup directional. Chapters 03 and 07 examine how to store the resulting keys and values more cheaply.
Grant Sanderson's video animates the lookup. It is a useful companion to the worked example above:
During training all positions are computed in parallelcontext An RNN must finish step t−1 before starting step t. A Transformer can train on all positions in a sequence at once, which suits GPUs. Chapter 04 returns to recurrence for inference., so position 5 could "see" position 9, which would let the model cheat at next-token prediction. The mask M sets every future-facing score to −∞context−∞ (a huge negative constant in practice) because e−∞ = 0: after the softmax, masked positions contribute nothing and the surviving weights renormalize to sum to 1 over the visible past. Masking after the softmax instead would leave a weight vector that no longer sums to one, a subtly broken average. before the softmax, forcing its weight to zero. Each token attends only backward. This triangle-shaped constraint is why decoder-only models can generate: the prediction at position t depends only on tokens 1…t.
A dot product of two random d-dimensional vectors with unit-variance entries has standard deviation ≈ √d. At d = 64 raw scores commonly reach magnitudes around 8 and beyond, and softmax of large scores saturates into a one-hot, killing gradients. Dividing by √dh keeps scores near ±1 so softmax stays in its responsive regime. (Introduced in Attention Is All You Need, Vaswani et al., 2017, §3.2.1.)
GPT-2 runs 12 of these lookups per layer, each with its own WQ, WK, WV over a 64-dim slicecontext One 768-dimensional head produces one set of weights for each token. Twelve smaller heads can produce twelve different sets. Their dimension, 64, also sets the state size discussed in chapter 04.. Different heads can learn different retrieval patterns, including syntax, coreference, and recency. Their outputs are concatenatedcontextConcatenated, then blended by one more learned projection (WO) that lets heads share findings. Heads turn out to be substantially redundant in trained models. GQA and MLA use some of that redundancy to reduce cache size (chapters 03 and 07).. Chapter 07 revisits how many key and value slots the model keeps.
Researchers have identified attention heads that help models complete patterns from a prompt without changing their weights. One example is an induction head.
What it does (the definition). Olsson et al. define an induction head by two
observable properties. Take a sequence containing "… Ms. Dursley … Ms." and ask what
follows the second Ms.:
Ms. it lands on Dursley (the token that had a Ms. in front of
it).Dursley's logit. Prediction: Dursley. The rule, in one line:
find where this context repeated before, and copy what came next.How it can work. To attend to a token whose predecessor matches
the current token, Dursley's key needs information about the preceding
Ms. One known mechanism uses two heads across two layers. First, an earlier
previous-token head copies each token's predecessor into it, writing "Ms.
came before" into Dursley's residual vector. Then the induction head's
query (from the second Ms.) matches that planted
key, and its value does the copying. Two simple heads,
can produce pattern completion.
This uses the query/key distinction from §2.1. The query seeks a token that followed a previous match, regardless of what the current token means on its own. Anthropic's team found these heads emerge reliably during training and argued they are a major part of in-context learning, a model's ability to pick up a pattern from the prompt itself at inference time, with no weight changes (Olsson et al., 2022). Fixed-size memories have trouble retaining the same ability to find an earlier occurrence and copy what followed it. That helps explain why the hybrid in chapter 07 still includes exact-attention layers.
An induction head has to reach back to an arbitrary earlier position and read what's there. That is cheap when every past token is still individually addressable (this chapter's attention) and difficult or impossible when history has been compressed into a fixed-size summary (chapter 04). Chapters 03–07 work through the cost of preserving this ability.
The example sentence below comes from Jay Alammar's explanation of how attention can connect a pronoun to an earlier noun:
Quoted from "The Illustrated Transformer" · Jay Alammar
"When the model is processing the word 'it', self-attention allows it to associate 'it' with 'animal'."
jalammar.github.io/illustrated-transformer, discussing the sentence used in the demo below.
One attention layer at context T does T² score computations at training time, and stores T keys + T values for generation. At T = 1,024 (GPT-2), that is manageable. At T = 1,048,576 (Kimi K3's context), T² exceeds a trillion pairs per head per layer. Chapter 03 calculates the memory cost; chapters 04–07 examine ways to reduce it.
What this chapter established
Go deeper
Contents · Glossary · Example sentence and quote credited to Jay Alammar; videos by 3Blue1Brown.