Chapter 01 · How a language model starts

The next-token machine

A language model predicts the next token from the text so far. That remains true for the frontier models of 2026. To see how their memory systems developed, start with an earlier decoder-only Transformer: GPT-2 small, released in 2019. It has 124 million parameters and one unmanaged form of exact context memory. Its architecture is also easy to inspect through several visual explanationscontext Bycroft's 3D model, the Transformer Explainer, and Karpathy's build-along all use this configuration. Its weights are public, so you can follow the same model across the chapter and those resources..

Memory contract · the ancestor (reference: GPT-2 small)
Stored
Learned weights (124M numbers, fixed after training) and temporary activations produced while processing the current context. The activations do not persist between separate predictions.
Growth
Fixed at training time. Context is capped at 1,024 tokens.
Eviction policy
Not applicable yet: there is no runtime memory to manage. That changes in chapter 03.
Failure mode
Anything beyond 1,024 tokens simply does not exist for the model.

1.1The training objective

A language model assigns probabilities to sequences. GPT-2 is autoregressive: it factorizes the probability of a text into a chain of next-token predictions, each conditioned on everything before it:

P(x1, …, xT) = P(x1) · P(x2 | x1) · P(x3 | x1, x2) ⋯ P(xT | x1, …, xT−1)

Training reduces the surprise (cross-entropy)context"Surprise" is −log P(true next token). If the model gives the actual token 1% probability, the score is −log(0.01) ≈ 4.6; at 90%, it is about 0.1. Lowering that score across many examples rewards the model for learning patterns in syntax, facts, style, and logic. of the actual next token across roughly 40 GB of web text (the WebText dataset). The objective does not specify a grammar module or a fact database. It asks the model to predict the next token.

The 3Blue1Brown video below gives a 27-minute visual introduction. The chapter also explains each part in text, so you can continue without watching it.

Transformers, the tech behind LLMs | Deep Learning Chapter 5 Video by 3Blue1Brown (Grant Sanderson) · embedded with attribution · interactive article version · watch on YouTube ↗

1.2Tokens: the model's alphabet

Models process tokens, chunks from a fixed vocabulary learned by byte-pair encodingcontextBPE builds its vocabulary by repeatedly merging the most frequent adjacent pair in the training text. The resulting vocabulary works like a compression codebook. Digits and rare names may split into awkward pieces, which can make tasks such as counting letters or long arithmetic harder.. GPT-2's vocabulary has 50,257 tokens. Common words can be single tokens; rare words may split into pieces (Kimi might become K+imi). This matters in two ways:

1.3Every token becomes a point in space

The model needs numbers to process cat. Its embedding table assigns that token a vector, a list of 768 numbers. You can treat the vector as a point in 768-dimensional space. Training tends to place tokens used in similar ways near one another, and some directions in that space carry meaning. A familiar example is that the direction from "man" to "woman" resembles the direction from "king" to "queen." Later layers work with these vectors rather than with words directly.

The vector for bites starts the same wherever the token appears. Attention also needs information about word order, so GPT-2 adds a learned position vectorcontextAttention itself is order-blind (a bag of tokens looks the same from every position), so position must be injected explicitly. GPT-2 learned one vector per position slot; modern models mostly rotate query/key pairs instead (RoPE), which generalizes better to long contexts. In either case, the model needs position information to distinguish different word orders. to each token's vector. You can think of it as adding a seat number. "Dog bites man" and "man bites dog" then enter the model as different inputs.

For a visual explanation, see the Embedding section of 3Blue1Brown's article or video. Alammar's Illustrated GPT-2 shows the token's full path through the model.

1.4What happens inside a block

GPT-2 small has 12 blocks. Each repeats the same two kinds of computation, using its own learned weights:

A prompt moves through the model in this order:

  1. Your text becomes tokens (§1.2); each token becomes vector + seat number (§1.3).
  2. Block 1: attention, then the MLP. Each result is added to the token's running vector.
  3. Blocks 2 through 12: the same two steps, with different learned weights. Earlier blocks tend to handle grammatical patterns; later ones handle more abstract patterns.
  4. The final vector at the last position exits to the head (§1.5), which turns it into a prediction.
Plate 1·AOne GPT-2 block: the only moving part, stacked twelve deep
residual stream (768 dims per token) causal self-attention the conversation step + MLP (768 → 3072 → 768) the quiet-thinking step + from previous block to next block · every branch output is ADDED to the stream LN LN
The small LN boxes are layer normalization, housekeeping that rescales vectors before each step; safe to ignore until chapter 09. The vertical spine is the residual stream, a term from Anthropic's A Mathematical Framework for Transformer Circuits (Elhage et al., 2021). Keep your eye on it: chapter 09 is entirely about this one line.

The + circles in the figure show that each step's output is added onto the streamcontextAdding instead of replacing (residual connections, from 2015's ResNets) is why very deep networks train at all: the identity path gives gradients an unobstructed highway, and each layer only has to learn a small correction. Chapter 09 examines what this design costs at 93 layers.. Each token's vector therefore carries a running total across the blocks. Chapter 09 returns to that accumulation.

The Attention Block and Multilayer Perceptron sections of 3Blue1Brown's article animate these steps. The Karpathy track in the Lab builds the block in code.

1.5The head: from a vector back to words

After 12 blocks, the last token's vector contains information for predicting the next token. The head turns its 768 numbers into probabilities in two steps. First, compare the vector against every vocabulary token's direction (using the same tied embedding table from §1.2), producing 50,257 match scores, the logits. Second, softmax turns those raw scores into a proper probability distribution: exponentiate each score, then divide by the total, so everything is positive and sums to 1. The result assigns a probability to every vocabulary token.

To generate text, the model samples from the distributioncontextHow you sample is a user-facing dial: divide the logits by a "temperature" before softmax. Low temperature sharpens toward the single best token (deterministic, sometimes repetitive); higher flattens the distribution (diverse, riskier). Chapter 02's demo uses the same softmax calculation for attention weights. at the last position. It appends that token and runs the model again. Every model in this explainer generates text through that repeated predict-sample-append process.

The Unembedding and Softmax sections of 3Blue1Brown's article show both steps.

1.6GPT-2 small, by the numbers

QuantityValueWhy it matters later
Total parameters124MThe denominator of the famous 22,580× ratio (ch. 10)
Layers (blocks)12K3 has 93 attention layers, so depth becomes its own memory problem (ch. 09)
Hidden size768 (12 heads × 64; heads are ch. 02's parallel mini-attentions)Head size sets the size of linear attention's state matrix (ch. 04)
Context window1,024 tokensvs. 1,048,576 in K3, a 1,024× stretch that drives chapters 03–07
Vocabulary50,257 BPE tokensK3 uses 160K, so vocabularies grew too

Figures: GPT-2 paper (Radford et al., 2019) and The Illustrated GPT-2; K3 figures from the Kimi K3 model card, verified 2026-08-15.

1.7Where those 124 million numbers come from

The 124 million weights begin as random values. Before training, GPT-2 produces gibberish. Training adjusts those values through a repeated process:

  1. Take a chunk of real text from the 40 GB pile. Hide the next token.
  2. Run the machine and read off its predicted distribution over the vocabulary.
  3. Score the surprise: how little probability it put on the token that actually came next.
  4. Adjust the weights in the direction that would have lowered that surprise, using backpropagationcontextBackpropagation is the chain rule from calculus, run in reverse through the network. For each weight, it computes how a small increase would change the prediction error. Those changes form the gradient. Gradient descent then steps the weights downhill. (In practice this runs on minibatches of examples at once, and an optimizer averages many such gradients before each step; most individual nudges are tiny.) Karpathy's first Zero-to-Hero lecture builds this from scratch in ~100 lines; see the Lab..
  5. Repeat across billions of examples from the corpus.

Each adjustment is small. Across many examples, weights that lower prediction error tend to capture patterns in grammar, facts, style, and reasoning. Many different weight settings can work; there is no single correct set. No one separately programmed a grammar module. For the rest of this explainer, the model is trained and its weights stay fixed during use.

1.8Why prediction can teach more than words

Good prediction compresses information about the text. To predict well, a model can benefit from representing the situations the text describes.

Consider a murder mystery. To predict the final sentence, "the killer was ___," a model that merely knew word frequencies would guess a common name and lose. A model that had actually tracked the alibis, the timeline, and the inconsistency in chapter 7 assigns high probability to the right name. That kind of prediction rewards a model for tracking information across the story. It does not, by itself, prove that the model understands the story as a person would.

This is why capabilities nobody separately supervised (translation, arithmetic, code, chain-of-thought reasoning, all present in the raw text but never singled out as training targets) can appear as models scale. The examples are present in the training text, even though the objective does not name them as separate tasks. The following chapters examine the memory systems that make prediction practical to compute.

Limit: prediction does not verify facts

The objective rewards likely continuations. It does not check whether a claim is true, so fluent answers can contain invented facts. Alignment training, including RLHF, addresses some of this behavior but is outside this explainer's scope.

Source spotlight: see this chapter in 3D and in the browser
Explore the model in 3D or run it in your browser
Brendan Bycroft's LLM Visualization renders a GPT-2-class model as an explorable 3D structure, every weight matrix of this chapter, spatially laid out. Transformer Explainer (Georgia Tech's Polo Club) runs a live GPT-2 in your browser: type text, watch real attention weights and probabilities form.

What this chapter established

Go deeper

Contents · Glossary · Sources credited inline; verified 2026-08-15.