Chapter 01 · How a language model starts
A language model predicts the next token from the text so far. That remains true for the frontier models of 2026. To see how their memory systems developed, start with an earlier decoder-only Transformer: GPT-2 small, released in 2019. It has 124 million parameters and one unmanaged form of exact context memory. Its architecture is also easy to inspect through several visual explanationscontext Bycroft's 3D model, the Transformer Explainer, and Karpathy's build-along all use this configuration. Its weights are public, so you can follow the same model across the chapter and those resources..
A language model assigns probabilities to sequences. GPT-2 is autoregressive: it factorizes the probability of a text into a chain of next-token predictions, each conditioned on everything before it:
Training reduces the surprise (cross-entropy)context"Surprise" is −log P(true next token). If the model gives the actual token 1% probability, the score is −log(0.01) ≈ 4.6; at 90%, it is about 0.1. Lowering that score across many examples rewards the model for learning patterns in syntax, facts, style, and logic. of the actual next token across roughly 40 GB of web text (the WebText dataset). The objective does not specify a grammar module or a fact database. It asks the model to predict the next token.
The 3Blue1Brown video below gives a 27-minute visual introduction. The chapter also explains each part in text, so you can continue without watching it.
Models process tokens,
chunks from a fixed vocabulary learned by
byte-pair encodingcontextBPE
builds its vocabulary by repeatedly merging the most frequent adjacent pair in the training text.
The resulting vocabulary works like a compression codebook. Digits and rare names may split into
awkward pieces, which can make tasks such as counting letters or long arithmetic harder..
GPT-2's vocabulary has 50,257 tokens. Common words can be single tokens; rare words may split
into pieces (Kimi might become K+imi). This matters in two ways:
The model needs numbers to process cat. Its embedding table assigns that token a
vector, a list of 768 numbers. You can treat the vector as a point in
768-dimensional space. Training tends to place tokens used in similar ways near one another,
and some directions in that space carry meaning. A familiar example is that the direction
from "man" to "woman" resembles the direction from "king" to "queen." Later layers work with
these vectors rather than with words directly.
The vector for bites starts the same wherever the token appears. Attention also
needs information about word order, so GPT-2 adds a learned
position vectorcontextAttention
itself is order-blind (a bag of tokens looks the same from every position), so position must be
injected explicitly. GPT-2 learned one vector per position slot; modern models mostly rotate
query/key pairs instead (RoPE), which generalizes better to long contexts. In either case,
the model needs position information to distinguish different word orders. to each
token's vector. You can think of it as adding a seat number. "Dog bites man" and "man bites dog"
then enter the model as different inputs.
For a visual explanation, see the Embedding section of 3Blue1Brown's article or video. Alammar's Illustrated GPT-2 shows the token's full path through the model.
GPT-2 small has 12 blocks. Each repeats the same two kinds of computation, using its own learned weights:
it can use information from
animal. Chapter 02 works through the calculation. GPT-2 performs it in 12 parallel
channels called heads.A prompt moves through the model in this order:
The + circles in the figure show that each step's output is
added onto the streamcontextAdding
instead of replacing (residual connections, from 2015's ResNets) is why very deep networks train at
all: the identity path gives gradients an unobstructed highway, and each layer only has to learn a
small correction. Chapter 09 examines what this design costs at 93 layers..
Each token's vector therefore carries a running total across the blocks. Chapter 09 returns
to that accumulation.
The Attention Block and Multilayer Perceptron sections of 3Blue1Brown's article animate these steps. The Karpathy track in the Lab builds the block in code.
After 12 blocks, the last token's vector contains information for predicting the next token. The head turns its 768 numbers into probabilities in two steps. First, compare the vector against every vocabulary token's direction (using the same tied embedding table from §1.2), producing 50,257 match scores, the logits. Second, softmax turns those raw scores into a proper probability distribution: exponentiate each score, then divide by the total, so everything is positive and sums to 1. The result assigns a probability to every vocabulary token.
To generate text, the model samples from the distributioncontextHow you sample is a user-facing dial: divide the logits by a "temperature" before softmax. Low temperature sharpens toward the single best token (deterministic, sometimes repetitive); higher flattens the distribution (diverse, riskier). Chapter 02's demo uses the same softmax calculation for attention weights. at the last position. It appends that token and runs the model again. Every model in this explainer generates text through that repeated predict-sample-append process.
The Unembedding and Softmax sections of 3Blue1Brown's article show both steps.
| Quantity | Value | Why it matters later |
|---|---|---|
| Total parameters | 124M | The denominator of the famous 22,580× ratio (ch. 10) |
| Layers (blocks) | 12 | K3 has 93 attention layers, so depth becomes its own memory problem (ch. 09) |
| Hidden size | 768 (12 heads × 64; heads are ch. 02's parallel mini-attentions) | Head size sets the size of linear attention's state matrix (ch. 04) |
| Context window | 1,024 tokens | vs. 1,048,576 in K3, a 1,024× stretch that drives chapters 03–07 |
| Vocabulary | 50,257 BPE tokens | K3 uses 160K, so vocabularies grew too |
Figures: GPT-2 paper (Radford et al., 2019) and The Illustrated GPT-2; K3 figures from the Kimi K3 model card, verified 2026-08-15.
The 124 million weights begin as random values. Before training, GPT-2 produces gibberish. Training adjusts those values through a repeated process:
Each adjustment is small. Across many examples, weights that lower prediction error tend to capture patterns in grammar, facts, style, and reasoning. Many different weight settings can work; there is no single correct set. No one separately programmed a grammar module. For the rest of this explainer, the model is trained and its weights stay fixed during use.
Good prediction compresses information about the text. To predict well, a model can benefit from representing the situations the text describes.
Consider a murder mystery. To predict the final sentence, "the killer was ___," a model that merely knew word frequencies would guess a common name and lose. A model that had actually tracked the alibis, the timeline, and the inconsistency in chapter 7 assigns high probability to the right name. That kind of prediction rewards a model for tracking information across the story. It does not, by itself, prove that the model understands the story as a person would.
This is why capabilities nobody separately supervised (translation, arithmetic, code, chain-of-thought reasoning, all present in the raw text but never singled out as training targets) can appear as models scale. The examples are present in the training text, even though the objective does not name them as separate tasks. The following chapters examine the memory systems that make prediction practical to compute.
The objective rewards likely continuations. It does not check whether a claim is true, so fluent answers can contain invented facts. Alignment training, including RLHF, addresses some of this behavior but is outside this explainer's scope.
What this chapter established
Go deeper
Contents · Glossary · Sources credited inline; verified 2026-08-15.