Chapter 07 · Reading the literature

The moves papers make

Open a 2026 architecture paper and the wall of notation suggests that research runs on math you never learned. Mostly, it doesn't. A surprisingly small repertoire of moves, applied with taste, powers most of what ships: approximate the big matrix with skinny ones (low-rank), keep the magnitudes civil (normalize), spend fewer bits per number (quantize). This chapter teaches the repertoire, one worked example each, then hands you a notation survival kit. The goal: papers stop reading as walls and start reading as familiar moves in new positions.

Where this lives in a transformer · chapter 07
The idea
Big learned matrices are usually wasteful: their useful content concentrates in far fewer directions than they have. Most efficiency research is one bet, repeated: find the waste, spend less on it.
Where you'll meet it
LoRA fine-tuning, DeepSeek's MLA, GQA, quantized inference (fp8, int4, GGUF), and the abstracts of a large fraction of arXiv's daily ML feed.
The one picture
A (big × big) matrix replaced by (big × r)·(r × big) with r tiny: the low-rank sandwich. Once you see it, you see it everywhere.
If you remember one thing
When a paper says "low-rank," read: "we bet the useful part of this matrix is a few directions, and kept only those." When it says a new number format, read: "we found precision nobody was using."

7.1The low-rank move, and the theorem behind it

Chapter 02 said every matrix is a function. Here is the refinement that unlocks half of efficiency research: every matrix is a function that can be split into rotate, stretch, rotate. That is the singular value decomposition (SVD): M = UΣVᵀ, with U and V pure rotations (plus possible reflections) and Σ a diagonal stretch. All the "strength" lives in Σ's diagonal entries, the singular values: sorted largest first, they say how much the function amplifies its 1st, 2nd, 3rd… most important input direction.

The empirical fact that turns this from algebra into money: for real-world matrices, the singular values collapse fast. A photo, a similarity table, a learned weight update: typically a handful of large values, then a long tail of near-zeros, meaning the matrix mostly acts through a few directions and barely does anything with the rest. So truncate: keep the top r directions, discard the tail, and store the result as two skinny matrices. The Eckart–Young theoremsoapboxFormally: truncating the SVD to its top r singular values gives the best possible rank-r approximation of the matrix (in Frobenius/spectral norm). "Best" is doing real work: no cleverer rank-r scheme can beat the lazy-sounding "keep the biggest stretches." Proved in 1936; earning industrial keep since at least the 1990s (search engines, recommender systems), and lately the math under cheap fine-tuning. guarantees this lazy-sounding recipe is optimal: no rank-r approximation does better. Sixty seconds with Tim Baumann's SVD image-compression slider will burn the intuition in: drag r up from 1 and watch a photo assemble from its dominant directions.

Plate 7·AThe move itself: one fat matrix ≈ two skinny ones
W 4096 × 4096 16.8M numbers ≈ B 4096 × 16 · A 16 × 4096 together: 131K numbers · 0.8% of W the bottleneck r = 16 every input is squeezed through 16 dimensions: only 16 directions of behavior survive. The bet is that 16 is all the task ever needed.
Shape-check it, chapter 02 style: (4096×16)·(16×4096) → (4096×4096). Same interface as W, fraction of the storage and compute. The inner dimension r is the rank of the product, and choosing r is choosing how much nuance to afford.

7.2Worked example one: LoRA

The purest industrial use of the move is LoRA (low-rank adaptation, Hu et al., 2021), the standard way to fine-tune a big model on a budget. Full fine-tuning updates every weight: for one 4096×4096 matrix, 16.8M adjustable numbers, times hundreds of matrices, and you must store that whole delta for every task you tune. LoRA's bet: the change a fine-tune needs is low-rank. Adapting a general model to a specific task shifts behavior along a few directions, not along thousands. So freeze W entirely and learn only a rank-r correction:

W′ = W (frozen) + B·A   with r ≈ 4–64

B starts at zero, so training begins exactly at the base model and learns only the correction. Gradients (ch. 05) flow only into A and B.

The arithmetic is startling (mathbox below): typically under one percent of the parameters, with quality close to full fine-tuning on most tasks, and adapters swap like plugins because W never moved. When quality does fall short, that too is informative: it means the task's change genuinely needed more directions than r allowed. The bet is honest either way, and "increase r" is the dial.

7.3Worked example two: shrinking attention's memory

The same move, aimed at chapter 06. Serving long contexts means caching every token's key and value vectors (the "KV cache"), and at frontier scale that cache, not the weights, becomes the thing that won't fit. The field's answers are a family portrait of one idea. MQA/GQA: let many query heads share one key/value head (or a few groups): fewer distinct K/V's to store. MLA (DeepSeek, 2024): the full low-rank sandwich: store, per token, only a compressed latent vector (the bottleneck of Plate 7·A), and reconstruct keys and values from it on the fly: a 90%+ cache reduction in DeepSeek-V2's configuration. You can now read that sentence as math rather than marketing: the K and V projections were mostly redundant directions; MLA keeps the directions that matter and re-expands on demand. Welch Labs' animation of exactly this progression (MHA → MQA → MLA) is the best sixteen minutes this chapter can recommend, and the natural bridge to reading DeepSeek's actual papers.

7.4Numbers on a diet: precision as a budget

The other great lever never touches the matrices' shapes, only how many bits each number gets. A float is sign + exponent + mantissa (chapter's last new concept, honest): the exponent sets the range of magnitudes you can express, the mantissa sets the precision within that range. Every format is one decision about splitting a bit budget between them:

Plate 7·BFour ways to slice a number, drawn to scale
fp32training's traditional gold standard · 4 bytes 32 bits ±exponent 8 · rangemantissa 23 · precision fp16half the bytes; narrow range bites (overflow at 65,504) 16 bits exp 5mantissa 10 bf16same bytes, range-first: fp32's exponent, less precision 16 bits exp 8 (= fp32)mantissa 7 fp8 (e4m3)one byte per number · today's inference frontier, with per-block scaling factors 8 bits exp 4man 3
The bf16 story is the one to remember: training kept diverging in fp16 because gradients (ch. 05) span wild magnitudes and the 5-bit exponent overflowed. bf16's fix: sacrifice precision, keep fp32's range. Deep learning wants range more than precision, an empirical discovery now baked into every TPU and GPU. Flip these bits yourself at float.exposed.

Quantization pushes the same logic below training: for inference, weights can ride formats as skinny as 4 bits (int4, GGUF's k-quant family, and mixed-precision recipes like MXFP4) with modest quality loss, because a trained matrix's numbers are redundant the same way its directions were in §7.1. Halve the bits and you halve the memory and roughly halve the time spent hauling weights through the memory bus, which is what bounds generation speed when tokens come out one at a time. Same lens as everything else in this chapter: find precision nobody was using, spend it.

7.5The notation survival kit

Last tool: the symbols themselves. The field's notation is smaller than it looks; this table plus chapter 02's shape-checking habit decodes most equations you will actually meet:

SymbolRead it asChapter that taught it
x ∈ ℝᵈ"x is a d-dimensional real vector": announcing a shape01, 02
Wx + ba learned linear function, plus a learned offset (bias)02
‖x‖, ‖x‖₂the vector's length (norm)01
Wᵀ, xᵀytranspose: shape plumbing; xᵀy is just the dot product x·y01, 02
⊙elementwise multiply (Hadamard): gate one vector by another, entry by entry02 §2.5's spirit
σ(·), gelu(·)an elementwise nonlinearity: the kink between matrices02 §2.5
x ~ p, 𝔼[·]"x is drawn from distribution p"; 𝔼 is the average over such draws03, 04
argmaxi zᵢthe index of the biggest entry (softmax's hard-edged cousin)03
log p(x | θ), 𝓛log-likelihood; 𝓛 is a loss. −(1/N)Σ log p is chapter 04, always04
∇θ 𝓛, ∂𝓛/∂wthe gradient: every weight's sensitivity dial, stacked05
DKL(p ‖ q)extra surprise for believing q when truth is p; a leash in RLHF papers04
W ∈ ℝm×n, BAshape announcement for a matrix; adjacent capitals multiply (compose)02, 07
O(T²d)cost scaling: attention's bill from ch. 06, in the shapes' own language02, 06

And a skimming protocol that respects your time, offered as one practitioner's habit: read the abstract, then go straight to the architecture figure and the shapes. Ask of any scary equation: what are the shapes (that's the plumbing), where is the softmax (that's a budget being allocated), where is the loss (that's what the gradients serve)? Those three questions, which are chapters 02, 03, and 04–05 respectively, dissolve most walls. For everything else there's the Goodfellow notation table, the closest thing the field has to an official key.

The LoRA arithmetic, done honestly

One 4096×4096 attention projection: full fine-tuning trains 4096² ≈ 16.8M numbers. LoRA at r = 16 trains B (4096×16) + A (16×4096) = 2 × 65,536 ≈ 131K numbers: 0.78%.

Across a 7B-parameter model, applying LoRA to the attention projections at r = 16 typically trains a few tens of millions of parameters, under 1% of the model, which is why the fine-tune fits on a gaming GPU while the base model's quality carries over. The exact percentage moves with r and with which matrices you adapt; the order of magnitude is the point.

Storage bonus: shipping a task = shipping A and B (megabytes), not a full model copy (many gigabytes). QLoRA composes §7.4 with §7.2: base weights quantized to 4 bits, corrections trained in bf16 on top. Two moves, one sandwich.

7.6The best teachers for this

video series · first ~4 videos Singular Value Decomposition Steve Brunton · U. Washington Rotation-stretch-rotation, matrix approximation, and why truncation is optimal, from the best SVD teacher on video. Forty minutes covers everything §7.1 needs. interactive SVD Image Compression Demo Tim Baumann Drag the rank slider on any image (including your own) and feel low-rank approximation in ten seconds. The LoRA intuition lives here. article Parameter-Efficient Finetuning with LoRA Sebastian Raschka The W + BA picture, the intrinsic-dimension argument, honest experiments, and the parameter arithmetic, from a trusted explainer. video · 18 min How DeepSeek Rewrote the Transformer Welch Labs The KV cache, MHA → MQA → MLA, and low-rank compression in a real frontier architecture, beautifully animated. §7.3 in motion, and a genuine bridge to reading papers. visual guide · 60+ figures A Visual Guide to Quantization Maarten Grootendorst Bit layouts for every format in Plate 7·B and onward through GPTQ and GGUF. Far and away the best number-formats-for-ML explainer.

7.7Graduation: where to take this next

You now hold the working set: geometry for meaning, matrices as learned functions, softmax budgets, surprise as the objective, gradients as the learner, the block that assembles them, and the efficiency moves played on top. Two natural next steps, in rough order:

What you should now believe

Go deeper

Contents · Glossary · The syllabus · Sources credited inline; links verified 2026-08-07.