Chapter 01 · The geometry of meaning
Every word a language model reads is converted, before anything else happens, into a vector: a list of a few hundred to a few thousand numbers, which you can equally picture as a point in a space with that many dimensions. Every mechanism in the rest of this primer (attention, loss, gradients, the whole transformer) is machinery for moving those points around. So the first intuition to rebuild is the oldest one: meaning as position, and similarity as geometry.
You met vectors twice in school, and the two versions probably never merged. Physics said: a
vector is an arrow, a thing with length and direction. Programming said: a vector
is an array, float x[768], a plain list of numbers. Both are right,
and the whole trick of this chapter is holding both pictures at once. The list
[3, 2] is the arrow that goes 3 steps east and 2 steps north; the arrow is
just the list drawn on graph paper. Add a third number and the arrow lives in 3-D space. Add 765
more and it lives in a 768-dimensional space, which nobody can draw, but which behaves, for
everything we need, exactly like the spaces you can.
One habit to install immediately, because every paper assumes it: a vector's lengthsoapboxAlso called the norm, written ‖v‖. For our purposes it's the Pythagorean theorem run in n dimensions: square every entry, add them up, take the square root. Papers write it ‖v‖₂ when they're being careful (there are other norms; chapter 07 mentions why you'd care). and its direction are separate facts. Two vectors can point the same way while one is ten times longer. Much of the machinery ahead (cosine similarity below, LayerNorm in chapter 06) exists precisely to manage that distinction: usually the direction carries the meaning and the length is bookkeeping.
Here is the single most load-bearing operation in this primer. Take two vectors of the same size. Multiply them entry by entry, and add up the results:
Left side: what the computer does (multiply, add). Right side: what it means (θ is the angle between the arrows).
The computational reading is almost insultingly simple, a loop any programmer writes in one line. The geometric reading is where the payoff lives: that same multiply-and-add equals the product of the two lengths times the cosine of the angle between the arrows. Which means the dot product is an agreement meter:
There is an even more physical reading: a·b asks "how much of a's shadow falls along b?", scaled by b's own length (projection, in the textbook vocabulary). A solar panel produces the most power facing the sun straight on, less at a slant, none edge-on. That falloff is the cosine, and every attention score computed inside an LLM is exactly this: one vector asking how much another vector faces it.
One refinement you will meet constantly: cosine similarity is the dot product with the lengths divided out, leaving pure direction-agreement on a fixed −1 to 1 scale. Use it when you care only about "same topic?" and not "how long are these vectors?"; that is why it, rather than the raw dot product, is the standard metric in embedding search and RAG systems.
Take the law of cosines on the triangle formed by a, b, and the connecting side a − b: ‖a − b‖² = ‖a‖² + ‖b‖² − 2‖a‖‖b‖cos θ.
Now expand the left side using entries. ‖a − b‖² = Σ(aᵢ − bᵢ)² = Σaᵢ² + Σbᵢ² − 2Σaᵢbᵢ = ‖a‖² + ‖b‖² − 2(a·b).
Set the two expansions equal, cancel the shared terms, divide by −2: a·b = ‖a‖‖b‖cos θ. The geometry was hiding inside the algebra all along, and nothing about the argument cared how many dimensions there were: it works identically in 768.
Now the payoff. A language model's first layer is a giant lookup table, the embedding table: one learned vector per vocabulary token. GPT-2 small's table is 50,257 rows (its vocabulary) by 768 columns (its vector size): about 38 million numbers, learned, not designed. "Learned" here means (chapter 05 makes this precise) that training nudged these numbers, billions of times, in whatever direction made the model better at predicting text.
The remarkable, genuinely earned discovery is what those nudges converge to: tokens
that behave similarly end up near each other. Not because anyone told the model
"cat resembles dog," but because words used in interchangeable contexts get pushed toward
interchangeable positions; that's the only way to predict them interchangeably. And the structure
runs deeper than clustering. Some directions in the space carry meaning by themselves.
The classic demonstration, from the word2vec team
(Mikolov et al., 2013): take the vector for
king, subtract man, add woman, and the nearest vocabulary
vector to where you land (setting aside the three words you started from, as these demos do) is
queen. Vector arithmetic performing analogy.
King − man + woman ≈ queen genuinely works in word2vec-style embeddings, and directions for
tense, plurality, and country→capital genuinely exist. But the trick is cherry-picked in most
demos (it fails for plenty of word pairs), and inside a modern transformer the picture is
messier: after the first layer, a token's vector is progressively rewritten by context (chapter
06), so "the vector for bank" stops being one fixed point and becomes
riverbank or Chase depending on the sentence. Static maps are the right
starter intuition, not the full story.
Two engineering footnotes worth filing. First, what gets embedded is not words but tokenssoapboxChunks from a fixed vocabulary, learned by byte-pair encoding: common words are one token, rare words shatter into pieces. Two minutes of pasting text into Tiktokenizer will dissolve the "models read words" misconception for good.. Second, the same trick now embeds everything: sentences, images, code, songs. Any data you can map to vectors such that "similar things land nearby" buys you search, clustering, and recommendation for free. That is the entire premise of vector databases and retrieval-augmented generation, and it is this chapter's math doing all the work.
This primer's job is connective tissue; the individual ideas have world-class teachers already. For this chapter, in recommended order:
A fair objection at this point: "you keep drawing 2-D pictures and asserting they scale." Mostly they do; the dot product's algebra is dimension-blind. But high-dimensional spaces have one property worth knowing about because LLMs exploit it, and because your 3-D intuition actively denies it: there is far more room than you think.
In 2-D you can draw exactly 2 perpendicular directions; in 768-D, exactly 768. But relax the requirement from "exactly 90°" to "within a degree or two of 90°," and the count stops being 768 and becomes exponentially large: astronomically many directions, all nearly unrelated to one another. A useful mental picture for why: in high dimensions, two arrows built by flipping coins for each coordinate almost always land nearly perpendicular, because their dot product is a sum of hundreds of random ± terms that mostly cancel. Near-orthogonality stops being a special arrangement and becomes the default.
This is why a model whose vectors have only 768 slots can still keep tens of thousands of concepts distinguishable: concepts don't need a private axis each, just a direction sufficiently unlike the others. (Interpretability researchers call the resulting packing superpositionsoapboxThe working hypothesis, developed in Anthropic's interpretability work, that models store many more "features" than they have dimensions by assigning them nearly-orthogonal directions and tolerating the small crosstalk. 3Blue1Brown's "How might LLMs store facts" closes with a lovely demonstration of the counting argument..) When chapter 06 shows attention comparing vectors by dot product, remember this section: near-zero dot products are cheap and plentiful, so "these two tokens are unrelated" is easy for the geometry to express.
What you should now believe
Go deeper
Contents · Glossary · The syllabus · Sources credited inline; links verified 2026-08-07.