Chapter 05 · The learning loop

Gradients: how models learn

Chapters 01–04 built a machine that turns text into a belief and a scoreboard that prices the belief's wrongness. One question remains, and it is the question: who sets the weights? GPT-2 has 124 million of them; frontier models, trillions. Nobody sets them. They start as noise, and a three-idea loop (the derivative as a sensitivity dial, the gradient as a compass, the chain rule as a blame-routing system) turns that noise into competence. This chapter rebuilds the loop; no derivation tables required.

Where this lives in a transformer · chapter 05
The idea
The derivative of the loss with respect to a weight answers one question: "if I nudge this weight up a hair, does my surprise rise or fall, and how fast?" Learning is asking that of every weight and stepping the answer.
Where you'll meet it
All of training. Also the vocabulary around it: optimizers (SGD, Adam), learning-rate schedules, "vanishing gradients," "gradient clipping," and why architectures are designed to be differentiable end to end.
The one equation
w ← w − η · ∂L/∂w : nudge every weight against its own sensitivity, at step size η.
If you remember one thing
Backpropagation is not a learning algorithm with opinions; it is the chain rule, mechanized, computing all million sensitivities for roughly the price of two forward passes.

5.1The derivative, rehabilitated

Forget the limit definition and the table of rules; here is the working meaning. The derivative of f at a point answers: if I nudge the input by a tiny amount, what happens to the output? It is a local exchange rate. ∂f/∂x = 3 means "a hair more x buys three hairs more f, right here" (elsewhere the rate may differ). ∂f/∂x = −0.5 means nudging x up pushes f down at half strength. Zero means f is locally indifferent to x: flat ground.

That's the entire concept, and notice it is an engineer's concept: a sensitivity you could measure empirically by wiggling the input and watching the output (finite differences, the numerical check every autograd test suite uses). Calculus just computes the wiggle-response without doing the wiggling. When the function has millions of inputs, as a loss over weights does, each weight gets its own partial derivative: its private answer to "does nudging me help or hurt?" while everyone else holds still.

5.2The gradient: every sensitivity at once, pointing uphill

Collect all those partial derivatives into one vector and you have the gradient, ∇L: a list with one entry per weight (for GPT-2, a 124-million-entry list), which chapter 01 taught you to also read as an arrow. The geometric fact that makes everything work: the gradient points in the direction of steepest increase of the loss, and its negation points steepest downhill. So the learning rule writes itself: compute the surprise (chapter 04), compute the gradient, take a small step the other way, repeat:

w ← w − η ∇L(w)

η (eta) is the learning rate: the step size. The most consequential hyperparameter in deep learning is this one scalar.

Plate 5·ALoss as landscape; learning as descent
low loss each ring: equal loss, like elevation lines start: random weights, high surprise each arrow: −η∇L, downhill by the local compass, recomputed after every step ∇L points uphill; we go the other way
Two honest caveats about this picture. The real landscape has 124 million axes, not two. And it is nothing like a smooth bowl: it is astronomically bumpy, and the goal is not the minimum (no one finds that) but any of the vast number of low, flat regions that predict text well. That such regions are reliably findable by local downhill steps is an empirical miracle the field leans on daily.
Walk downhill yourselfinteractive · step the ball

One weight, one loss curve, the update rule w ← w − η·L′(w). Try η ≈ 0.1 (crawls), η ≈ 0.5 (efficient), η ≈ 1.0 (overshoots and oscillates), η at the max (diverges). You have now experienced every learning-rate pathology in the literature.

5.3The chain rule: blame, routed backward

One puzzle stands between the update rule and reality. The loss depends on a weight in layer 2 only through a long relay: that weight nudges its layer's output, which nudges layer 3's input, which nudges layer 4's… all the way to the final surprise. How do you get one number, ∂L/∂w, out of a relay?

The chain rule, and it says something you already believe about relays: sensitivities multiply along a chain. If x moves y at rate 3, and y moves z at rate 2, then x moves z at rate 6. Gearboxes compose exactly this way; so do exchange rates (dollars→euros→yen). Where two paths from the same input reconverge, their contributions add. Multiply along paths, add across paths: that is the entire rule, and Christopher Olah's Calculus on Computational Graphs earns its classic status by showing how far those two verbs go.

Plate 5·BMultiply along the path, add where paths merge
weight wlayer 2 activation apath 1 activation bpath 2 logit zreconverged loss Lsurprise ×2.0 ×0.5 ×(−1.0) ×3.0 ×0.4 ∂L/∂w = path 1 + path 2 = (2.0 × −1.0 × 0.4) + (0.5 × 3.0 × 0.4) = −0.8 + 0.6 = −0.2 every edge label is a LOCAL derivative: cheap, known in closed form for each primitive op
The final −0.2 says: nudging w up lowers the loss slightly, so training will nudge it up. Multiply along paths, add across paths, and note that no step required anything smarter than arithmetic on local sensitivities.

5.4Backprop: the chain rule at industrial scale

Now the miracle, and it is a miracle of cost, not of concept. You need ∂L/∂w for every one of 124 million weights. Done naively (wiggle each weight, rerun the model), that is 124 million forward passes per learning step. The fix is an ordering trick worthy of a systems engineer: work backward from the loss, computing ∂L/∂(each node) layer by layer. Each node's sensitivity is built from the sensitivities of the nodes after it (which are already done) times cheap local derivatives. One backward sweep, touching each edge once, yields every weight's gradient simultaneously, at roughly twice one forward pass's cost, no matter how many weights there are. Backpropagation is exactly this: the chain rule plus dynamic programming. Without the trick, training a trillion-parameter model is not expensive; it is impossible. With it, "differentiate the whole model" became a button (loss.backward()), which is why every framework is built around an autograd engine and why architectures are designed end-to-end differentiable: the button has to reach everything.

Soapbox: why engineers should still look under this hood

Karpathy's essay "Yes you should understand backprop" makes the case with production scars: backprop is a leaky abstraction. Sigmoids saturate and their local ×(near-zero) silently zeroes every upstream gradient; ReLUs "die"; RNN gradients explode through repeated multiplication. Each pathology is obvious if you picture the multiply-along-paths relay, and invisible if you only ever pressed the button. The fix for one of them (make the relay additive, so gradients flow through untouched) is the residual connection, load-bearing in chapter 06.

5.5The realities: batches, optimizers, schedules

Three refinements separate the textbook loop from what actually runs on the cluster, and all three are one sentence each. Stochastic gradient descent: computing the true gradient over the whole corpus per step is absurd, so each step uses a random minibatch: a noisy but unbiased compass, traded billions of cheap noisy steps for millions of perfect ones. Momentum and Adam: raw SGD oscillates across steep ravines while crawling along gentle floors, so practical optimizers keep a running memory of recent gradients (momentum) and normalize each weight's step by the typical size of its gradients (Adam, the LLM default; the interactive Distill article Why Momentum Really Works lets you feel the ravine problem directly). Schedules: η itself is steered over training, warmed up, then decayed, because early noise and late fine-tuning want different step sizes. (And when a rogue batch produces a monster gradient, it is clipped to a maximum length before the step: gradient clipping, the seatbelt you'll see in every training config.) None of this changes the story; all of it is chasing the same downhill compass more efficiently.

The prettiest gradient in deep learning: softmax + cross-entropy

Chapters 03 and 04 composed softmax (logits z → probabilities q) with cross-entropy loss (L = −log qcorrect). Differentiate the composite with respect to the logits and an avalanche of chain-rule terms collapses to:

∂L/∂zᵢ = qᵢ − yᵢ  (y is 1 at the true token, 0 elsewhere)

The gradient at the output is literally prediction minus truth. Predict 0.70 for the right token: its logit feels −0.30 (push up). Predict 0.05 for a wrong one: +0.05 (push down, mildly). The learning signal is the error itself, clean, bounded, and never saturating. This tidy pairing is a real reason the softmax/cross-entropy combination, rather than a plausible alternative, became universal: the field keeps objectives whose gradients are kind. (Full derivation: CS231n's linear-classification notes; it is four honest lines of quotient rule.)

5.6The best teachers for this

video · 21 min Gradient descent, how neural networks learn (DL Ch. 2) 3Blue1Brown The canonical animation of §5.1–§5.2: loss landscapes, the downhill compass, and what "learning" looks like as geometry. video · 13 min Backpropagation, intuitively (DL Ch. 3) 3Blue1Brown Blame flowing backward through a network, animated. Watch it after §5.3 and the chain rule stops being algebra. course notes CS231n: Backpropagation, Intuitions Stanford · originally Andrej Karpathy The most engineer-friendly written backprop ever: circuits with local gradients, and the mnemonic "add distributes, max routes, multiply swaps and scales." interactive TensorFlow Playground Smilkov & Carter · Google Train a real (tiny) network in your browser: watch the loss fall, the weights thicken, and the learning rate matter. A decade old, still unbeaten.

What you should now believe

Go deeper

Contents · Glossary · The syllabus · Sources credited inline; links verified 2026-08-07.