Chapter 02 · Functions in disguise
Strip away the schedulers, the kernels, and the marketing, and an LLM spends almost all of its arithmetic doing one thing: matrix multiplication. If matrices are, in your memory, "grids of numbers you multiply by an unmemorable rule," this chapter is the repair job. The working intuition: a matrix is a function, one that takes a vector in and gives a vector out, and the unmemorable rule is just that function being applied. Once that clicks, a transformer stops looking like a wall of algebra and starts looking like what it is: a pipeline of learned functions.
Start with the smallest possible claim and build. A 2×2 matrix times a 2-vector gives another 2-vector. So the matrix maps vectors to vectors: input in, output out. That makes it a function. But a very special kind, a linear one: it keeps grid lines parallel and evenly spaced, and it leaves the origin where it is. No bending, no curving, no shifting. Every linear function on vectors, in any dimension, is exactly one matrix, and every matrix is exactly one such function. The grid of numbers is just the function's storage format.
And the format is wonderfully legible, once you know where to look: the columns of the matrix are where the basis vectors land. In 2-D: column one is where the unit arrow pointing east ends up; column two is where the north arrow ends up. Every other vector's fate follows by linearity. Read a matrix column by column and you are reading the function's behavior directly.
Now jump to LLM scale, where the honest headline is: a matrix is a function you can learn. A 768×768 weight matrix in GPT-2 is a learned function from 768-dimensional vectors to 768-dimensional vectors, with 589,824 adjustable numbers determining what it does. Training (chapter 05) tunes those numbers. Nobody designs the rotation; the gradient finds it. When a paper says "we add a learned projection," translate: "we insert an adjustable linear function here and let training decide what it should do."
Here is the rule you memorized for (matrix × vector), reframed so it's obvious. To compute Mv: each output entry is a dot product between one row of M and the input vector. That's the whole rule. And after chapter 01, you can read meaning into it: every row of the matrix is a little probe asking "how much does the input agree with this pattern?" A matrix is a bank of dot-product questions; the output vector is the list of answers.
There is a second, equally true reading that pays off constantly: Mv is a mix of the columns of M, using v's entries as the recipe (v₁ of column 1, plus v₂ of column 2, …). Rows ask questions; columns get mixed. Both are the same arithmetic, and fluent readers of papers flip between them without noticing. Matrix × matrix is then nothing new: to compute AB, apply A to each column of B separately. Which yields the fact that makes deep learning possible to write down at all:
Read right to left, always: B happens to v first, then A happens to the result.
"Deep" in deep learning is this line, taken seriously: stack many learned functions so the composite can be rich even though each stage is simple. One caution to carry along: AB ≠ BA in general. Order matters, "rotate then stretch" and "stretch then rotate" are different functions, and half the subscript discipline in papers exists to keep the order straight.
Engineers already own the perfect metaphor for the last rule of this chapter: matrix shapes are a type system. A matrix of shape (m×n) is a function from n-vectors to m-vectors; it simply does not accept other sizes. Chaining matrices type-checks like chaining functions: (m×n)·(n×p) works, the inner n's must match, and they cancel to leave (m×p). When you read a paper, running this check in your head is how you decode what every symbol is: find the shapes and the equation explains itself.
Run the check on one real example, the attention scores you will meet properly in chapter 06. A sentence of T tokens, each a 768-vector, is stored as an (T×768) matrix X (one row per token: a spreadsheet, tokens down, features across). One bookkeeping flip to notice before the table: with tokens as rows, the learned matrix now multiplies on the right (X·W), where §2.1's pictures had it on the left (Mv). Same functions, transposed layout; papers use both freely, and the shapes always tell you which one you're reading. Then:
| Step | Shapes | What just happened, in words |
|---|---|---|
| Q = X·WQ | (T×768)·(768×64) → (T×64) | every token's vector is pushed through the same learned function, producing a 64-dim "what am I looking for?" vector per token |
| K = X·WK | (T×768)·(768×64) → (T×64) | same trick, different learned function: a "what do I offer?" vector per token |
| scores = Q·KT | (T×64)·(64×T) → (T×T) | every query dot-producted against every key: a full table of token-to-token agreement, chapter 01's meter run T² times |
Notice what the shape-check just bought you: without knowing anything else about attention, you already know the result is T×T, one number per pair of tokens, and you know why long contexts get expensive (double T and the table quadruples). That transposesoapboxThe T flips a matrix over its diagonal: rows become columns. Here it's pure shape plumbing, turning K into the orientation where the dot products line up. When you see a transpose in a paper, the author is almost always just making the types check. on K is there purely to make the types line up. Papers are full of moves like this; they look profound and are plumbing.
Matrix multiplication is not just conceptually central; it is economically central. GPUs and TPUs are, to first approximation, matmul appliances (NVIDIA literally ships "tensor cores"), and their entire design bet is that one operation dominates. It's why "how many FLOPs does this model need" is answerable from shapes alone, and why architecture research (chapter 07) is substantially the art of getting the same intelligence out of smaller multiplies.
One loose end, and it is load-bearing. §2.2 celebrated composition, but composition has a trapdoor: a stack of purely linear functions collapses. If every layer is a matrix, the whole network is (A·B·C⋯), which is… one matrix. A hundred linear layers have exactly the expressive power of a single one. All that depth, wasted.
The fix is almost embarrassingly cheap: between matrices, apply a simple nonlinear function to each entry of the vector separately. The classic is ReLU: replace negative entries with zero, keep the rest. (GPT-2 uses GELUsoapboxA smoothed ReLU (the corner at zero is rounded off). Modern models mostly use gated variants like SwiGLU. The differences matter for training stability and a percent or two of quality; for intuition, "smooth kink" covers all of them., a smoothed version.) That kink is enough. With it, stacking stops collapsing, and depth buys genuinely new functions: the composite can bend, gate, and carve up space in ways no single matrix can. The MLP inside every transformer block is exactly this sandwich:
Two learned matrices and one fixed kink: that's the "quiet thinking" half of every block, holding two-thirds of each block's parameters.
So the division of labor in all of deep learning is: matrices do the mixing, nonlinearities make the depth count. Everything else is arrangement.
A function f is linear when f(u + v) = f(u) + f(v) and f(cv) = c·f(v). Compose two of them: g(f(u + v)) = g(f(u) + f(v)) = g(f(u)) + g(f(v)), and likewise for scaling. So g∘f passes the same two tests, is itself linear, and is therefore representable as a single matrix. Induct: any depth of linear layers is one matrix. The nonlinearity between layers is what breaks the induction, and it is the only thing in the stack that does.
What you should now believe
Go deeper
Contents · Glossary · The syllabus · Sources credited inline; links verified 2026-08-07.