A guided refresher · seven chapters · curated from the best teachers on the internet

The Math Behind LLMs and Transformers

Rebuild the intuition, not the coursework: the five mathematical moves under every transformer

Varma Chanderraju · August 2026

You are a software engineer. Every AI research paper, every architecture card, every training log you encounter assumes you know and remember the math underpinning LLMs. Here is the encouraging secret this primer is built on: the mathematical core of a transformer is five moves: the dot product as a similarity meter, the matrix as a learned function, softmax as a budget-maker, cross-entropy as a surprise bill, and the gradient as a downhill compass. Rebuild honest intuition for those five and the wall of notation becomes a set of familiar parts. That rebuild, visual and unhurried, is the whole job of the next seven chapters.

Plate 0·AThe five moves, and where each one lives inside a transformer
DOT PRODUCT similarity meter meaning as geometry ch. 01 MATRIX a learned function shapes = types ch. 02 SOFTMAX scores → budget of belief ch. 03 SURPRISE −log p · the loss every curve plots ch. 04 GRADIENT downhill compass ch. 05 THE TRANSFORMER BLOCK attention + MLP on a residual stream: the five moves, arranged · ch. 06 READING PAPERS WITHOUT FLINCHING low-rank · quantization · the notation survival kit · ch. 07
The "five moves" framing is this primer's own connective tissue; every mechanism under it belongs to the sources credited inline throughout.

What is this?

A refresher and a guided tour, not a course. The target reader is someone who has encountered these concepts before, but the math has gone quiet from disuse. The goal is the working mental model: enough intuition that when you dig into training or inference internals, open a model card, or hit an equation in a paper, the parts feel recognizable and the notation decodes. No problem sets, no proofs for their own sake (the few derivations worth seeing sit in collapsible boxes you may freely skip), and no pretense of mastery. Chapters are meant to be read front to back; each one leans on the previous.

Authorship: this primer is a curation as much as a composition. The individual ideas here have some of the best teachers alive (3Blue1Brown, Steve Brunton, Andrej Karpathy, Jay Alammar, Chris Olah, and a dozen more), and it would be malpractice to paraphrase them at lower quality. So each chapter builds the intuition in prose and original figures, then hands you cards pointing at the best teaching of that topic on the internet, chosen from a verified shortlist and annotated with why each earns your time. The syllabus page collects all of it as one watch-and-read path.

How to read it

Each chapter opens with a small ledger stating where its math lives inside a transformer and the one thing to remember. Four colors mean the same thing in every figure and equation:

vectors & directions: where meaning lives matrices & transformations: learned functions probabilities & distributions: budgets of belief gradients & change: how learning moves

Dotted-underlined words like softmax link into the glossary (hover for the definition). Dashed-underlined phrases like this onesoapboxHover or tap a dashed phrase anywhere in the primer and a small aside like this bubbles out: extra context, an opinion, or a why, without breaking the paragraph you're in. are inline asides. A few figures are interactive (drag the arrows in chapter 01, the temperature slider in chapter 03, the descent stepper in chapter 05); everything else that moves lives at the external links, where the people who built those interactives did it better than a static page could.

The chapters

Chapter 01Vectors and meaningArrow and array are one object; the dot product measures agreement; embeddings turn words into geometry. Drag two arrows and feel it. Chapter 02Matrices as transformationsA matrix is a learned function; multiplying is composing; shapes are a type system. Why depth needs a kink. Chapter 03From scores to probabilitiesLogits, softmax, and the temperature dial, with a live distribution to bend. The model's entire interface is one budget. Chapter 04Loss, surprise, and informationSurprise = −log p, entropy is the floor, cross-entropy is the bill, perplexity is the humane rescale. Every loss curve, decoded. Chapter 05Gradients: how models learnDerivative as sensitivity dial, gradient as compass, backprop as the chain rule at industrial scale. Walk a loss curve yourself. Chapter 06The transformer block, assembledThe attention equation read symbol by symbol with chapter numbers, then heads, the residual stream, LayerNorm, and RoPE. Chapter 07The moves papers makeLow-rank (SVD → LoRA → MLA), number formats (fp32 → fp8), and a notation survival kit. Graduation: read the real thing. The syllabusWatch & readEvery curated resource in one place, as a sequenced path: what to watch, what to read, and why each one earned its slot. ReferenceGlossaryEvery term, defined once, linked from everywhere.

Source materials

Every recommendation card says what the resource teaches and why it beat its alternatives. The core teachers, beyond papers linked where used: 3Blue1Brown (linear algebra and the neural-network series), Steve Brunton (SVD), Andrej Karpathy (backprop and GPT from scratch), Jay Alammar (the illustrated series), Chris Olah (information theory, backprop), Georgia Tech's Polo Club and Brendan Bycroft (interactive transformers), EleutherAI (the RoPE explainer; the technique itself is Su et al.'s), Sebastian Raschka (LoRA, sampling, attention variants), and Maarten Grootendorst (quantization). Where a chapter's figure adapts someone's idea, the caption says so.

Soapbox: why another math explainer at all?

Because the great ones each teach an island: 3Blue1Brown the linear algebra, Olah the information theory, Karpathy the backprop. What a rusty engineer can't easily get is the through-line: one sequenced path, in one vocabulary, where the dot product you rebuild on Monday is visibly the same operation inside the attention equation you read on Friday, and where every hop ends at the best existing teacher rather than a paraphrase of them. The through-line is this primer's contribution; the teaching it points to belongs to the people it names.

Scoped claim: what this primer is not

Not a substitute for a linear algebra course, not sufficient preparation for research-level theory, and not a survey of architectures. If you want the math done properly, with proofs, the free references on the syllabus page (Mathematics for Machine Learning; Dive into Deep Learning) are the honest deep end.

Created by Varma Chanderraju. Built with Claude, Codex and Gemini.