Chapter 07 · Reading the literature
Open a 2026 architecture paper and the wall of notation suggests that research runs on math you never learned. Mostly, it doesn't. A surprisingly small repertoire of moves, applied with taste, powers most of what ships: approximate the big matrix with skinny ones (low-rank), keep the magnitudes civil (normalize), spend fewer bits per number (quantize). This chapter teaches the repertoire, one worked example each, then hands you a notation survival kit. The goal: papers stop reading as walls and start reading as familiar moves in new positions.
Chapter 02 said every matrix is a function. Here is the refinement that unlocks half of efficiency research: every matrix is a function that can be split into rotate, stretch, rotate. That is the singular value decomposition (SVD): M = UΣVᵀ, with U and V pure rotations (plus possible reflections) and Σ a diagonal stretch. All the "strength" lives in Σ's diagonal entries, the singular values: sorted largest first, they say how much the function amplifies its 1st, 2nd, 3rd… most important input direction.
The empirical fact that turns this from algebra into money: for real-world matrices, the singular values collapse fast. A photo, a similarity table, a learned weight update: typically a handful of large values, then a long tail of near-zeros, meaning the matrix mostly acts through a few directions and barely does anything with the rest. So truncate: keep the top r directions, discard the tail, and store the result as two skinny matrices. The Eckart–Young theoremsoapboxFormally: truncating the SVD to its top r singular values gives the best possible rank-r approximation of the matrix (in Frobenius/spectral norm). "Best" is doing real work: no cleverer rank-r scheme can beat the lazy-sounding "keep the biggest stretches." Proved in 1936; earning industrial keep since at least the 1990s (search engines, recommender systems), and lately the math under cheap fine-tuning. guarantees this lazy-sounding recipe is optimal: no rank-r approximation does better. Sixty seconds with Tim Baumann's SVD image-compression slider will burn the intuition in: drag r up from 1 and watch a photo assemble from its dominant directions.
The purest industrial use of the move is LoRA (low-rank adaptation, Hu et al., 2021), the standard way to fine-tune a big model on a budget. Full fine-tuning updates every weight: for one 4096×4096 matrix, 16.8M adjustable numbers, times hundreds of matrices, and you must store that whole delta for every task you tune. LoRA's bet: the change a fine-tune needs is low-rank. Adapting a general model to a specific task shifts behavior along a few directions, not along thousands. So freeze W entirely and learn only a rank-r correction:
B starts at zero, so training begins exactly at the base model and learns only the correction. Gradients (ch. 05) flow only into A and B.
The arithmetic is startling (mathbox below): typically under one percent of the parameters, with quality close to full fine-tuning on most tasks, and adapters swap like plugins because W never moved. When quality does fall short, that too is informative: it means the task's change genuinely needed more directions than r allowed. The bet is honest either way, and "increase r" is the dial.
The same move, aimed at chapter 06. Serving long contexts means caching every token's key and value vectors (the "KV cache"), and at frontier scale that cache, not the weights, becomes the thing that won't fit. The field's answers are a family portrait of one idea. MQA/GQA: let many query heads share one key/value head (or a few groups): fewer distinct K/V's to store. MLA (DeepSeek, 2024): the full low-rank sandwich: store, per token, only a compressed latent vector (the bottleneck of Plate 7·A), and reconstruct keys and values from it on the fly: a 90%+ cache reduction in DeepSeek-V2's configuration. You can now read that sentence as math rather than marketing: the K and V projections were mostly redundant directions; MLA keeps the directions that matter and re-expands on demand. Welch Labs' animation of exactly this progression (MHA → MQA → MLA) is the best sixteen minutes this chapter can recommend, and the natural bridge to reading DeepSeek's actual papers.
The other great lever never touches the matrices' shapes, only how many bits each number gets. A float is sign + exponent + mantissa (chapter's last new concept, honest): the exponent sets the range of magnitudes you can express, the mantissa sets the precision within that range. Every format is one decision about splitting a bit budget between them:
Quantization pushes the same logic below training: for inference, weights can ride formats as skinny as 4 bits (int4, GGUF's k-quant family, and mixed-precision recipes like MXFP4) with modest quality loss, because a trained matrix's numbers are redundant the same way its directions were in §7.1. Halve the bits and you halve the memory and roughly halve the time spent hauling weights through the memory bus, which is what bounds generation speed when tokens come out one at a time. Same lens as everything else in this chapter: find precision nobody was using, spend it.
Last tool: the symbols themselves. The field's notation is smaller than it looks; this table plus chapter 02's shape-checking habit decodes most equations you will actually meet:
| Symbol | Read it as | Chapter that taught it |
|---|---|---|
| x ∈ ℝᵈ | "x is a d-dimensional real vector": announcing a shape | 01, 02 |
| Wx + b | a learned linear function, plus a learned offset (bias) | 02 |
| ‖x‖, ‖x‖₂ | the vector's length (norm) | 01 |
| Wᵀ, xᵀy | transpose: shape plumbing; xᵀy is just the dot product x·y | 01, 02 |
| ⊙ | elementwise multiply (Hadamard): gate one vector by another, entry by entry | 02 §2.5's spirit |
| σ(·), gelu(·) | an elementwise nonlinearity: the kink between matrices | 02 §2.5 |
| x ~ p, 𝔼[·] | "x is drawn from distribution p"; 𝔼 is the average over such draws | 03, 04 |
| argmaxi zᵢ | the index of the biggest entry (softmax's hard-edged cousin) | 03 |
| log p(x | θ), 𝓛 | log-likelihood; 𝓛 is a loss. −(1/N)Σ log p is chapter 04, always | 04 |
| ∇θ 𝓛, ∂𝓛/∂w | the gradient: every weight's sensitivity dial, stacked | 05 |
| DKL(p ‖ q) | extra surprise for believing q when truth is p; a leash in RLHF papers | 04 |
| W ∈ ℝm×n, BA | shape announcement for a matrix; adjacent capitals multiply (compose) | 02, 07 |
| O(T²d) | cost scaling: attention's bill from ch. 06, in the shapes' own language | 02, 06 |
And a skimming protocol that respects your time, offered as one practitioner's habit: read the abstract, then go straight to the architecture figure and the shapes. Ask of any scary equation: what are the shapes (that's the plumbing), where is the softmax (that's a budget being allocated), where is the loss (that's what the gradients serve)? Those three questions, which are chapters 02, 03, and 04–05 respectively, dissolve most walls. For everything else there's the Goodfellow notation table, the closest thing the field has to an official key.
One 4096×4096 attention projection: full fine-tuning trains 4096² ≈ 16.8M numbers. LoRA at r = 16 trains B (4096×16) + A (16×4096) = 2 × 65,536 ≈ 131K numbers: 0.78%.
Across a 7B-parameter model, applying LoRA to the attention projections at r = 16 typically trains a few tens of millions of parameters, under 1% of the model, which is why the fine-tune fits on a gaming GPU while the base model's quality carries over. The exact percentage moves with r and with which matrices you adapt; the order of magnitude is the point.
Storage bonus: shipping a task = shipping A and B (megabytes), not a full model copy (many gigabytes). QLoRA composes §7.4 with §7.2: base weights quantized to 4 bits, corrections trained in bf16 on top. Two moves, one sandwich.
You now hold the working set: geometry for meaning, matrices as learned functions, softmax budgets, surprise as the objective, gradients as the learner, the block that assembles them, and the efficiency moves played on top. Two natural next steps, in rough order:
What you should now believe
Go deeper
Contents · Glossary · The syllabus · Sources credited inline; links verified 2026-08-07.