Chapter 01 · The geometry of meaning

Vectors and meaning

Every word a language model reads is converted, before anything else happens, into a vector: a list of a few hundred to a few thousand numbers, which you can equally picture as a point in a space with that many dimensions. Every mechanism in the rest of this primer (attention, loss, gradients, the whole transformer) is machinery for moving those points around. So the first intuition to rebuild is the oldest one: meaning as position, and similarity as geometry.

Where this lives in a transformer · chapter 01
The idea
A token's meaning is its position in a high-dimensional space. Related meanings sit near each other; some directions in the space carry meaning of their own.
Where you'll meet it
The embedding table that turns tokens into vectors, the query·key match scores inside attention (chapter 06), and the final layer that scores every vocabulary word against the model's output vector.
The one operation
The dot product: a·b = a₁b₁ + a₂b₂ + ⋯ + aₙbₙ. Multiply matching entries, add everything up, get one number.
If you remember one thing
Vectors that point the same way mean similar things, and the dot product is the machine's one-number measure of "how much do these point the same way?"

1.1A vector is two pictures of the same thing

You met vectors twice in school, and the two versions probably never merged. Physics said: a vector is an arrow, a thing with length and direction. Programming said: a vector is an array, float x[768], a plain list of numbers. Both are right, and the whole trick of this chapter is holding both pictures at once. The list [3, 2] is the arrow that goes 3 steps east and 2 steps north; the arrow is just the list drawn on graph paper. Add a third number and the arrow lives in 3-D space. Add 765 more and it lives in a 768-dimensional space, which nobody can draw, but which behaves, for everything we need, exactly like the spaces you can.

Plate 1·AOne object, two readings: the list is the arrow
the engineer's picture v = [ 3, 2 ] an array of floats index 0 → 3 index 1 → 2 = the physicist's picture 3 steps east 2 steps north v length and direction, drawn from the origin
Every vector in this primer is both at once: a list you could print, and a direction something is pointing. When a sentence below says "points the same way," you may always translate it back to "has a similar pattern of numbers."

One habit to install immediately, because every paper assumes it: a vector's lengthsoapboxAlso called the norm, written ‖v‖. For our purposes it's the Pythagorean theorem run in n dimensions: square every entry, add them up, take the square root. Papers write it ‖v‖₂ when they're being careful (there are other norms; chapter 07 mentions why you'd care). and its direction are separate facts. Two vectors can point the same way while one is ten times longer. Much of the machinery ahead (cosine similarity below, LayerNorm in chapter 06) exists precisely to manage that distinction: usually the direction carries the meaning and the length is bookkeeping.

1.2The dot product: multiply-and-add measures agreement

Here is the single most load-bearing operation in this primer. Take two vectors of the same size. Multiply them entry by entry, and add up the results:

a · b  =  a₁b₁ + a₂b₂ + ⋯ + aₙbₙ  =  ‖a‖ ‖b‖ cos θ

Left side: what the computer does (multiply, add). Right side: what it means (θ is the angle between the arrows).

The computational reading is almost insultingly simple, a loop any programmer writes in one line. The geometric reading is where the payoff lives: that same multiply-and-add equals the product of the two lengths times the cosine of the angle between the arrows. Which means the dot product is an agreement meter:

There is an even more physical reading: a·b asks "how much of a's shadow falls along b?", scaled by b's own length (projection, in the textbook vocabulary). A solar panel produces the most power facing the sun straight on, less at a slant, none edge-on. That falloff is the cosine, and every attention score computed inside an LLM is exactly this: one vector asking how much another vector faces it.

Feel the dot productinteractive · drag the arrows
a b
a·b = …

Drag the arrowheads. Watch the sign flip as the angle passes 90°, and the magnitude swell when the arrows align or grow. Attention (chapter 06) computes millions of exactly these numbers per sentence.

One refinement you will meet constantly: cosine similarity is the dot product with the lengths divided out, leaving pure direction-agreement on a fixed −1 to 1 scale. Use it when you care only about "same topic?" and not "how long are these vectors?"; that is why it, rather than the raw dot product, is the standard metric in embedding search and RAG systems.

The derivation, if you want it: why multiply-and-add equals ‖a‖‖b‖cos θ

Take the law of cosines on the triangle formed by a, b, and the connecting side a − b: ‖a − b‖² = ‖a‖² + ‖b‖² − 2‖a‖‖b‖cos θ.

Now expand the left side using entries. ‖a − b‖² = Σ(aᵢ − bᵢ)² = Σaᵢ² + Σbᵢ² − 2Σaᵢbᵢ = ‖a‖² + ‖b‖² − 2(a·b).

Set the two expansions equal, cancel the shared terms, divide by −2: a·b = ‖a‖‖b‖cos θ. The geometry was hiding inside the algebra all along, and nothing about the argument cared how many dimensions there were: it works identically in 768.

1.3Words become points: the embedding table

Now the payoff. A language model's first layer is a giant lookup table, the embedding table: one learned vector per vocabulary token. GPT-2 small's table is 50,257 rows (its vocabulary) by 768 columns (its vector size): about 38 million numbers, learned, not designed. "Learned" here means (chapter 05 makes this precise) that training nudged these numbers, billions of times, in whatever direction made the model better at predicting text.

The remarkable, genuinely earned discovery is what those nudges converge to: tokens that behave similarly end up near each other. Not because anyone told the model "cat resembles dog," but because words used in interchangeable contexts get pushed toward interchangeable positions; that's the only way to predict them interchangeably. And the structure runs deeper than clustering. Some directions in the space carry meaning by themselves. The classic demonstration, from the word2vec team (Mikolov et al., 2013): take the vector for king, subtract man, add woman, and the nearest vocabulary vector to where you land (setting aside the three words you started from, as these demos do) is queen. Vector arithmetic performing analogy.

Plate 1·BA 2-D shadow of embedding space (real spaces: hundreds to thousands of dims)
animals cat dog kitten puppy finance bank loan interest king queen man woman the same "female direction," twice "royalty direction" Nearness = used in similar contexts. Parallel arrows = a direction that means something by itself (gender, tense, plurality, country→capital…).
Schematic, not real data: any 2-D drawing of a 768-dimensional space is a cartoon. For the real thing, open the TensorFlow Embedding Projector and fly through actual word2vec vectors in 3-D.
Scoped claim: analogies are real but fragile

King − man + woman ≈ queen genuinely works in word2vec-style embeddings, and directions for tense, plurality, and country→capital genuinely exist. But the trick is cherry-picked in most demos (it fails for plenty of word pairs), and inside a modern transformer the picture is messier: after the first layer, a token's vector is progressively rewritten by context (chapter 06), so "the vector for bank" stops being one fixed point and becomes riverbank or Chase depending on the sentence. Static maps are the right starter intuition, not the full story.

Two engineering footnotes worth filing. First, what gets embedded is not words but tokenssoapboxChunks from a fixed vocabulary, learned by byte-pair encoding: common words are one token, rare words shatter into pieces. Two minutes of pasting text into Tiktokenizer will dissolve the "models read words" misconception for good.. Second, the same trick now embeds everything: sentences, images, code, songs. Any data you can map to vectors such that "similar things land nearby" buys you search, clustering, and recommendation for free. That is the entire premise of vector databases and retrieval-augmented generation, and it is this chapter's math doing all the work.

1.4The best teachers for this

This primer's job is connective tissue; the individual ideas have world-class teachers already. For this chapter, in recommended order:

video · 27 min Transformers, the tech behind LLMs (Deep Learning Ch. 5) 3Blue1Brown · Grant Sanderson The embedding segment (roughly minutes 12–18) animates words-as-directions on real GPT weights, including the king/queen arithmetic. The best 6 minutes on this chapter's topic anywhere. visual essay The Illustrated Word2Vec Jay Alammar Builds embeddings from a "personality scores" analogy up to how they're trained. The canonical gentle on-ramp, all pictures, zero prerequisites. interactive Embedding Projector TensorFlow team · Google Fly through real embeddings in 3-D, click a word, see its nearest neighbors by cosine distance. The only way to feel the neighbor structure firsthand. article · short Understanding the Dot Product Kalid Azad · BetterExplained Dot product as "directional multiplication," with the solar-panel picture. If §1.2 didn't fully land, this will finish the job.

1.5High dimensions are not like 3-D, only bigger

A fair objection at this point: "you keep drawing 2-D pictures and asserting they scale." Mostly they do; the dot product's algebra is dimension-blind. But high-dimensional spaces have one property worth knowing about because LLMs exploit it, and because your 3-D intuition actively denies it: there is far more room than you think.

In 2-D you can draw exactly 2 perpendicular directions; in 768-D, exactly 768. But relax the requirement from "exactly 90°" to "within a degree or two of 90°," and the count stops being 768 and becomes exponentially large: astronomically many directions, all nearly unrelated to one another. A useful mental picture for why: in high dimensions, two arrows built by flipping coins for each coordinate almost always land nearly perpendicular, because their dot product is a sum of hundreds of random ± terms that mostly cancel. Near-orthogonality stops being a special arrangement and becomes the default.

This is why a model whose vectors have only 768 slots can still keep tens of thousands of concepts distinguishable: concepts don't need a private axis each, just a direction sufficiently unlike the others. (Interpretability researchers call the resulting packing superpositionsoapboxThe working hypothesis, developed in Anthropic's interpretability work, that models store many more "features" than they have dimensions by assigning them nearly-orthogonal directions and tolerating the small crosstalk. 3Blue1Brown's "How might LLMs store facts" closes with a lovely demonstration of the counting argument..) When chapter 06 shows attention comparing vectors by dot product, remember this section: near-zero dot products are cheap and plentiful, so "these two tokens are unrelated" is easy for the geometry to express.

What you should now believe

Go deeper

Contents · Glossary · The syllabus · Sources credited inline; links verified 2026-08-07.