Reference · every term, defined once
Glossary
Definitions favor the working intuition over the textbook formality; each entry
names the chapter that builds it properly. Dotted-underlined terms across the primer link here,
and hovering any of them shows these definitions in place.
- vector ch. 01
- A list of numbers and an arrow in space: the same object read two ways. In LLMs, the
universal container of meaning: every token, every intermediate state, every gradient is one.
Where it matters: everything downstream operates on vectors.
- dot product ch. 01
- Multiply two same-length vectors entry by entry and add it all up: one number measuring
directional agreement (positive = aligned, zero = unrelated, negative = opposed). Equals
‖a‖‖b‖cos θ.
Where it matters: every attention score is one.
- cosine similarity ch. 01
- The dot product with both lengths divided out: pure direction-agreement on a fixed −1 to 1
scale. The standard "same topic?" metric in embedding search and RAG.
Where it matters: vector databases, semantic search.
- norm ‖v‖ ch. 01
- A vector's length: square every entry, sum, take the root (the ₂ in ‖v‖₂ marks this
standard recipe among several). Direction usually carries the meaning; the norm is bookkeeping
that LayerNorm keeps civil.
Where it matters: LayerNorm, gradient clipping, paper notation.
- embedding ch. 01
- The learned lookup table (one vector per vocabulary token) and, by extension, any mapping
of data into vectors where nearness means similarity. Training pushes interchangeable tokens
toward interchangeable positions.
Where it matters: the model's first layer, and all of vector search.
- token ch. 01
- The unit models actually read: a chunk from a fixed vocabulary (learned by byte-pair
encoding in GPT-family models; other tokenizers exist), not a word. Common words are one token;
rare words shatter.
Where it matters: context lengths, pricing, and why models are odd at spelling.
- matrix ch. 02
- A stored linear function on vectors: its columns are where the input axes land. A learned
matrix is a learned function; a GPT is a few dozen of them with plumbing.
Where it matters: every weight you've ever heard of.
- matrix multiplication ch. 02
- Applying one linear function to another's output: multiplying matrices composes functions.
Row-times-column is a bank of dot products; shapes chain like types: (m×n)·(n×p) → (m×p).
Where it matters: ~all of an LLM's arithmetic; what GPUs are built for.
- rank / low-rank ch. 07
- The number of independent directions a matrix actually uses. Low-rank = the useful behavior
concentrates in few directions, so the matrix can be stored as two skinny factors with a small
inner dimension r.
Where it matters: LoRA, MLA, and half of efficiency research.
- nonlinearity (ReLU, GELU) ch. 02
- The cheap entrywise kink between matrices (ReLU zeroes negatives; GELU rounds the corner).
Without it, any stack of linear layers collapses into one matrix; with it, depth buys new
functions.
Where it matters: inside every MLP block.
- probability distribution ch. 03
- A budget of belief: one non-negative share per possible outcome, summing to exactly 1. An
LLM's output at every step is one, over its whole vocabulary.
Where it matters: the model's entire interface with the world.
- logits ch. 03
- The raw pre-softmax scores, one per vocabulary token, produced by dot-producting the final
vector against every token's direction. Only their differences mean anything.
Where it matters: what temperature rescales; what the loss differentiates.
- softmax ch. 03
- Exponentiate every score, divide by the total: scores in, honest probability budget out.
Rankings survive; gaps become ratios. A soft, differentiable "pick the max."
Where it matters: the output layer, and every attention row.
- temperature ch. 03
- Divide logits by T before softmax. Below 1 sharpens toward the favorite; above 1 flattens
toward uniform. One dial repricing every probability ratio at once, named for its Boltzmann
ancestry.
Where it matters: the sampling knob every API exposes.
- entropy ch. 04
- A distribution's expected surprise about its own draws: −Σ p log p. The built-in suspense,
and the floor no predictor of that source can average below.
Where it matters: why loss curves flatten; language's own entropy is the asymptote.
- cross-entropy ch. 04
- Average surprise when reality picks outcomes and your model q pays the bill: −Σ p log q.
For LLM training it collapses to −log q(actual next token). The number on every loss curve.
Where it matters: the pretraining objective of every LLM shipped.
- perplexity ch. 04
- e raised to the cross-entropy: the model's effective number of equally-likely choices per
token. Loss 2.3 ≈ "torn among 10 options." The humane rescaling of the same scoreboard.
Where it matters: evaluation tables in papers.
- KL divergence ch. 04
- Cross-entropy minus entropy: the extra surprise paid for believing q when the truth is p.
Zero when beliefs match; asymmetric otherwise.
Where it matters: the "leash" term in RLHF and distillation papers.
- loss function ch. 04
- The single number training minimizes; for LLMs, cross-entropy (average surprise at real
text). Everything the model becomes is a side effect of pushing this number down.
Where it matters: the objective the gradients serve.
- gradient ∇L ch. 05
- Every weight's sensitivity dial ("does nudging me raise or lower the loss, how fast?"),
stacked into one vector that points steepest uphill. Learning steps the other way.
Where it matters: the compass of all training.
- learning rate η ch. 05
- The step size in w ← w − η∇L. Too small crawls; too large oscillates or diverges. Warmed up
and decayed over training by a schedule.
Where it matters: the most consequential single hyperparameter.
- chain rule ch. 05
- Sensitivities multiply along a chain of dependencies (like gear ratios) and add where paths
merge. The entire mathematical content of backpropagation.
Where it matters: how blame reaches a layer-2 weight from the loss.
- backpropagation ch. 05
- The chain rule run backward from the loss with dynamic programming: every weight's gradient
in one sweep, at roughly twice a forward pass's cost. The button (
loss.backward())
that makes trillion-parameter training arithmetic instead of fantasy.
Where it matters: every autograd framework.
- attention ch. 06
- A soft, differentiable lookup: queries dot-producted against keys, softmaxed into budgets,
spent on a weighted average of values. softmax(QKᵀ/√d)V. Each token's new vector is a blend of
what it chose to read.
Where it matters: the mechanism that gives transformers context.
- query / key / value ch. 06
- Three learned linear projections of each token's vector: the question it asks (Q), the
advertisement it posts (K), the content it offers (V). Learned, not designed.
Where it matters: the vocabulary of every attention diagram.
- multi-head attention ch. 06
- Several small attentions run in parallel, each with its own Q/K/V matrices and its own
agenda (syntax, position, brackets…), concatenated and mixed by one more learned matrix. A
committee of specialist readers.
Where it matters: "12 heads × 64 dims" in every model card.
- residual stream ch. 06
- The running per-token vector that every block adds its output onto rather than
replacing. Simultaneously the gradients' unobstructed highway and the model's shared read/write
bus.
Where it matters: why depth trains at all; the load-bearing picture of interpretability.
- LayerNorm / RMSNorm ch. 06
- Recenter and rescale each token's vector to standard scale (LayerNorm), or skip the
recentering and just rescale (RMSNorm, the modern default), plus small learned dials, before
each block's branches. Directions carry meaning; the norm keeps a hundred layers of additions
from blowing the magnitudes out.
Where it matters: the "norm" boxes in every block diagram.
- RoPE (rotary position embedding) ch. 06
- Encode position by rotating each 2-D pair of query/key dimensions by an angle proportional
to the token's position. Dot products then depend only on relative distance, which
generalizes gracefully to long contexts.
Where it matters: how almost every modern LLM knows about order.
- SVD (singular value decomposition) ch. 07
- Every matrix split into rotate · stretch · rotate. The stretch factors (singular values),
sorted, reveal how much of the matrix's behavior lives in each direction; truncating to the top
r is the provably best rank-r approximation (Eckart–Young).
Where it matters: the theory under every low-rank trick.
- LoRA ch. 07
- Fine-tune by freezing W and learning a low-rank correction: W′ = W + BA with tiny inner
rank r. Under 1% of the parameters, adapters that swap like plugins, quality close to full
fine-tuning when the bet ("the change is low-rank") holds.
Where it matters: the default budget fine-tuning method.
- quantization ch. 07
- Spend fewer bits per number: fp32 → bf16 → fp8 → int4, trading precision (and sometimes
range) for memory and bandwidth. Trained weights tolerate starvation diets surprisingly well.
Where it matters: why a 70B model runs on your laptop at all.