Chapter 02 · Retrieval from earlier tokens

Attention: exact retrieval

Attention lets a token use information from earlier tokens. It compares the current token with each one, then combines their information according to how well they match. Because each earlier token has its own slot, the model can retrieve a specific part of the context. Chapters 04–07 examine ways to reduce the cost of that lookup.

Memory contract · softmax attention
Stored
One key ("what is stored here") and one value ("the content itself") vector per past token, per head; all three roles are defined properly in §2.1 just below
Growth
Linear in context length T; every token adds a slot
Eviction policy
None. Every past token stays individually addressable forever.
Failure mode
Cost: each new token's query must be compared against all T stored keys

2.1How the lookup works

Every token's vector is projected three ways, and the three projections play different roles. The same colors identify them throughout the explainer:

query: "what am I looking for?" key: "here's what I contain" value: "here's my actual content"
Q = XWQ   K = XWK   V = XWV

First, score every query against every key with a dot productcontext Dot products are efficient on GPUs: each dimension takes a multiply-add, and many scores can be computed as a matrix multiplication.. Next, softmaxcontext Softmax exponentiates the scores and divides by their sum. It increases the influence of better matches while keeping all weights positive and summing to 1. The slider below shows how changing score sharpness affects those weights. turns the scores into weights. Finally, use those weights to combine the values:

Attn(Q, K, V) = softmax(QK⊤√dh + M)V

M is the causal mask (§2.2); dh is the per-head dimension (64 in GPT-2).

In plain words, one attention step is five moves: ask (form a query), advertise (every past token exposes a key), score (dot products measure match), normalize (softmax turns scores into a probability-like distribution), blend (weighted sum of values).

The five moves, with actual numbers

Three tokens, the cat sat, and the query comes from sat. Suppose the scaled dot products of sat's query against the three keys came out to 2, 1, 0 (three plain numbers; how they arose is the projections' business). Softmax exponentiates and normalizes:

scores            2        1        0
e^score           7.389    2.718    1.000     (sum = 11.107)
weights           0.665    0.245    0.090     (each e^score / sum, they sum to 1)

Give the three tokens simple 2-dim values, v(the) = (10, 0), v(cat) = (0, 10), v(sat) = (4, 4), and blend:

output = 0.665·(10, 0) + 0.245·(0, 10) + 0.090·(4, 4)
       = (6.65, 0)    + (0, 2.45)     + (0.36, 0.36)
       = (7.01, 2.81)

The output at sat is (7.01, 2.81). The value from the has the largest weight, followed by cat and sat. A one-point difference between two raw scores produces about a 2.7× difference in weight. Real attention heads use learned projections and more dimensions, but the calculation is the same.

Why queries, keys, and values are separate

Each token needs three projections because looking for information, being found, and providing information are different jobs:

So the three projection matrices WQ, WK, WV are three learned questions asked of the same vector: what do I seek, how am I found, what do I give. Separate queries and keys make the lookup directional. Chapters 03 and 07 examine how to store the resulting keys and values more cheaply.

Plate 2·AThe lookup, drawn once: remember this shape
hidden states Xone vector / token Q K V QK⊤ / √dall-pairs scores mask + softmaxscores → weights × output = weights · V
The Q/K/V retrieval diagram follows The Illustrated Transformer by Jay Alammar, which explains the same steps visually.

Grant Sanderson's video animates the lookup. It is a useful companion to the worked example above:

Attention in transformers, step-by-step | Deep Learning Chapter 6 Video by 3Blue1Brown (Grant Sanderson) · embedded with attribution · interactive article version · watch on YouTube ↗

2.2The causal mask: no reading ahead

During training all positions are computed in parallelcontext An RNN must finish step t−1 before starting step t. A Transformer can train on all positions in a sequence at once, which suits GPUs. Chapter 04 returns to recurrence for inference., so position 5 could "see" position 9, which would let the model cheat at next-token prediction. The mask M sets every future-facing score to −∞context−∞ (a huge negative constant in practice) because e−∞ = 0: after the softmax, masked positions contribute nothing and the surviving weights renormalize to sum to 1 over the visible past. Masking after the softmax instead would leave a weight vector that no longer sums to one, a subtly broken average. before the softmax, forcing its weight to zero. Each token attends only backward. This triangle-shaped constraint is why decoder-only models can generate: the prediction at position t depends only on tokens 1…t.

Math checkpoint: why divide by √d?

A dot product of two random d-dimensional vectors with unit-variance entries has standard deviation ≈ √d. At d = 64 raw scores commonly reach magnitudes around 8 and beyond, and softmax of large scores saturates into a one-hot, killing gradients. Dividing by √dh keeps scores near ±1 so softmax stays in its responsive regime. (Introduced in Attention Is All You Need, Vaswani et al., 2017, §3.2.1.)

2.3Several lookups in parallel

GPT-2 runs 12 of these lookups per layer, each with its own WQ, WK, WV over a 64-dim slicecontext One 768-dimensional head produces one set of weights for each token. Twelve smaller heads can produce twelve different sets. Their dimension, 64, also sets the state size discussed in chapter 04.. Different heads can learn different retrieval patterns, including syntax, coreference, and recency. Their outputs are concatenatedcontextConcatenated, then blended by one more learned projection (WO) that lets heads share findings. Heads turn out to be substantially redundant in trained models. GQA and MLA use some of that redundancy to reduce cache size (chapters 03 and 07).. Chapter 07 revisits how many key and value slots the model keeps.

2.4An example: the induction head

Researchers have identified attention heads that help models complete patterns from a prompt without changing their weights. One example is an induction head.

What it does (the definition). Olsson et al. define an induction head by two observable properties. Take a sequence containing "… Ms. Dursley … Ms." and ask what follows the second Ms.:

  1. Prefix matching: the head attends from the current token back to an earlier token that was preceded by the same token we're at now, so from the second Ms. it lands on Dursley (the token that had a Ms. in front of it).
  2. Copying: the head then pushes the attended-to token's identity into the output, raising Dursley's logit. Prediction: Dursley. The rule, in one line: find where this context repeated before, and copy what came next.

How it can work. To attend to a token whose predecessor matches the current token, Dursley's key needs information about the preceding Ms. One known mechanism uses two heads across two layers. First, an earlier previous-token head copies each token's predecessor into it, writing "Ms. came before" into Dursley's residual vector. Then the induction head's query (from the second Ms.) matches that planted key, and its value does the copying. Two simple heads, can produce pattern completion.

This uses the query/key distinction from §2.1. The query seeks a token that followed a previous match, regardless of what the current token means on its own. Anthropic's team found these heads emerge reliably during training and argued they are a major part of in-context learning, a model's ability to pick up a pattern from the prompt itself at inference time, with no weight changes (Olsson et al., 2022). Fixed-size memories have trouble retaining the same ability to find an earlier occurrence and copy what followed it. That helps explain why the hybrid in chapter 07 still includes exact-attention layers.

Why exact lookup matters

An induction head has to reach back to an arbitrary earlier position and read what's there. That is cheap when every past token is still individually addressable (this chapter's attention) and difficult or impossible when history has been compressed into a fixed-size summary (chapter 04). Chapters 03–07 work through the cost of preserving this ability.

2.5Try an attention map

The example sentence below comes from Jay Alammar's explanation of how attention can connect a pronoun to an earlier noun:

Quoted from "The Illustrated Transformer" · Jay Alammar

"When the model is processing the word 'it', self-attention allows it to associate 'it' with 'animal'."

jalammar.github.io/illustrated-transformer, discussing the sentence used in the demo below.

Who is "it"? A causal attention row, liveInteractive

Click any token to make it the query. Earlier tokens shade amber by attention weight. The sharpness slider scales scores before softmax. Watch how softmax turns "slightly better match" into "winner takes most."

The raw match scores here are hand-authored for teaching, not taken from a trained model. The softmax is real. Two simplifications: the query token can't attend to itself here (real GPT masks allow self-attention), and there's no √d scaling. For genuine GPT-2 attention weights on your own text, use Transformer Explainer.

The cost of exact attention

One attention layer at context T does T² score computations at training time, and stores T keys + T values for generation. At T = 1,024 (GPT-2), that is manageable. At T = 1,048,576 (Kimi K3's context), T² exceeds a trillion pairs per head per layer. Chapter 03 calculates the memory cost; chapters 04–07 examine ways to reduce it.

What this chapter established

Go deeper

Contents · Glossary · Example sentence and quote credited to Jay Alammar; videos by 3Blue1Brown.