Chapter 05 · The learning loop
Chapters 01–04 built a machine that turns text into a belief and a scoreboard that prices the belief's wrongness. One question remains, and it is the question: who sets the weights? GPT-2 has 124 million of them; frontier models, trillions. Nobody sets them. They start as noise, and a three-idea loop (the derivative as a sensitivity dial, the gradient as a compass, the chain rule as a blame-routing system) turns that noise into competence. This chapter rebuilds the loop; no derivation tables required.
Forget the limit definition and the table of rules; here is the working meaning. The derivative of f at a point answers: if I nudge the input by a tiny amount, what happens to the output? It is a local exchange rate. ∂f/∂x = 3 means "a hair more x buys three hairs more f, right here" (elsewhere the rate may differ). ∂f/∂x = −0.5 means nudging x up pushes f down at half strength. Zero means f is locally indifferent to x: flat ground.
That's the entire concept, and notice it is an engineer's concept: a sensitivity you could measure empirically by wiggling the input and watching the output (finite differences, the numerical check every autograd test suite uses). Calculus just computes the wiggle-response without doing the wiggling. When the function has millions of inputs, as a loss over weights does, each weight gets its own partial derivative: its private answer to "does nudging me help or hurt?" while everyone else holds still.
Collect all those partial derivatives into one vector and you have the gradient, ∇L: a list with one entry per weight (for GPT-2, a 124-million-entry list), which chapter 01 taught you to also read as an arrow. The geometric fact that makes everything work: the gradient points in the direction of steepest increase of the loss, and its negation points steepest downhill. So the learning rule writes itself: compute the surprise (chapter 04), compute the gradient, take a small step the other way, repeat:
η (eta) is the learning rate: the step size. The most consequential hyperparameter in deep learning is this one scalar.
One puzzle stands between the update rule and reality. The loss depends on a weight in layer 2 only through a long relay: that weight nudges its layer's output, which nudges layer 3's input, which nudges layer 4's… all the way to the final surprise. How do you get one number, ∂L/∂w, out of a relay?
The chain rule, and it says something you already believe about relays: sensitivities multiply along a chain. If x moves y at rate 3, and y moves z at rate 2, then x moves z at rate 6. Gearboxes compose exactly this way; so do exchange rates (dollars→euros→yen). Where two paths from the same input reconverge, their contributions add. Multiply along paths, add across paths: that is the entire rule, and Christopher Olah's Calculus on Computational Graphs earns its classic status by showing how far those two verbs go.
Now the miracle, and it is a miracle of cost, not of concept. You need ∂L/∂w for
every one of 124 million weights. Done naively (wiggle each weight, rerun the model), that is 124
million forward passes per learning step. The fix is an ordering trick worthy of a systems
engineer: work backward from the loss, computing ∂L/∂(each node) layer by layer.
Each node's sensitivity is built from the sensitivities of the nodes after it (which are
already done) times cheap local derivatives. One backward sweep, touching each edge once, yields
every weight's gradient simultaneously, at roughly twice one forward pass's
cost, no matter how many weights there are.
Backpropagation is exactly this: the chain rule
plus dynamic programming. Without the trick, training a trillion-parameter model is not
expensive; it is impossible. With it, "differentiate the whole model" became a button
(loss.backward()), which is why every framework is built around an autograd engine
and why architectures are designed end-to-end differentiable: the button has to reach everything.
Karpathy's essay "Yes you should understand backprop" makes the case with production scars: backprop is a leaky abstraction. Sigmoids saturate and their local ×(near-zero) silently zeroes every upstream gradient; ReLUs "die"; RNN gradients explode through repeated multiplication. Each pathology is obvious if you picture the multiply-along-paths relay, and invisible if you only ever pressed the button. The fix for one of them (make the relay additive, so gradients flow through untouched) is the residual connection, load-bearing in chapter 06.
Three refinements separate the textbook loop from what actually runs on the cluster, and all three are one sentence each. Stochastic gradient descent: computing the true gradient over the whole corpus per step is absurd, so each step uses a random minibatch: a noisy but unbiased compass, traded billions of cheap noisy steps for millions of perfect ones. Momentum and Adam: raw SGD oscillates across steep ravines while crawling along gentle floors, so practical optimizers keep a running memory of recent gradients (momentum) and normalize each weight's step by the typical size of its gradients (Adam, the LLM default; the interactive Distill article Why Momentum Really Works lets you feel the ravine problem directly). Schedules: η itself is steered over training, warmed up, then decayed, because early noise and late fine-tuning want different step sizes. (And when a rogue batch produces a monster gradient, it is clipped to a maximum length before the step: gradient clipping, the seatbelt you'll see in every training config.) None of this changes the story; all of it is chasing the same downhill compass more efficiently.
Chapters 03 and 04 composed softmax (logits z → probabilities q) with cross-entropy loss (L = −log qcorrect). Differentiate the composite with respect to the logits and an avalanche of chain-rule terms collapses to:
∂L/∂zᵢ = qᵢ − yᵢ (y is 1 at the true token, 0 elsewhere)
The gradient at the output is literally prediction minus truth. Predict 0.70 for the right token: its logit feels −0.30 (push up). Predict 0.05 for a wrong one: +0.05 (push down, mildly). The learning signal is the error itself, clean, bounded, and never saturating. This tidy pairing is a real reason the softmax/cross-entropy combination, rather than a plausible alternative, became universal: the field keeps objectives whose gradients are kind. (Full derivation: CS231n's linear-classification notes; it is four honest lines of quotient rule.)
What you should now believe
Go deeper
Contents · Glossary · The syllabus · Sources credited inline; links verified 2026-08-07.