The syllabus · every curated resource, one sequenced path
Watch & read
Everything this primer recommends, gathered in one place and sequenced. Every link
was verified live on 2026-08-07; every entry says what the resource teaches and why it earned its
slot over the alternatives. Formats are tagged:
video,
interactive,
written.
S.1If you only have three hours: the spine
Five picks that carry the whole arc. Watch/play them in this order, reading the matching
chapters between sittings:
- Transformers, the tech behind LLMs · 3Blue1Brown · video, 27 min · after ch. 01. The full pipeline at a glance, with the embeddings-as-directions segment this primer's chapter 01 is built around.
- Attention in transformers, step-by-step · 3Blue1Brown · video, 26 min · after ch. 06. The finest visual explanation of attention in existence.
- Transformer Explainer · Polo Club, Georgia Tech · interactive, ~20 min of play · alongside ch. 03. A live GPT-2 in your browser: type, drag temperature, watch the distribution bend.
- Gradient descent, how neural networks learn · 3Blue1Brown · video, 21 min · after ch. 05. The canonical loss-landscape animation.
- LLM Visualization · Brendan Bycroft · interactive 3-D, ~30 min of exploration · after ch. 06. Fly through a working GPT; the architecture becomes a place you've been.
S.2Chapter 01 · Vectors and meaning
- The Illustrated Word2Vec · Jay Alammar · written. Embeddings from a "personality scores" analogy up through training; the canonical gentle on-ramp.
- Embedding Projector · TensorFlow team · interactive. Real embeddings in 3-D; click a word, see its cosine neighbors.
- Understanding the Dot Product · Kalid Azad · written, short. "Directional multiplication," with the solar-panel picture.
- Essence of Linear Algebra · 3Blue1Brown · video series, ~3 h total. The full rebuild; chapters 1, 2, and 9 (dot products) serve this chapter, chapters 3–4 serve the next.
- Embeddings: what they are and why they matter · Simon Willison · written. The engineer's tour: three lines of Python to search, RAG, and CLIP.
- Counterintuitive Properties of High Dimensional Space · Marc Khoury · written. Fifteen figures of 3-D intuition failing politely.
- How might LLMs store facts · 3Blue1Brown · video, 22 min. The superposition counting argument, animated (final third).
- Tiktokenizer · interactive. Two minutes that dissolve the "models read words" misconception.
S.3Chapter 02 · Matrices as transformations
- A Geometrical Understanding of Matrices · Gregory Gundersen · written. Columns-are-destinations, in essay form with figures.
- mm + Inside the Matrix · Basil Hosmer · interactive 3-D + written tour. Matmul as a navigable cube, with real GPT-2 attention-head presets.
- Immersive Linear Algebra · Ström, Åström & Akenine-Möller · interactive textbook. Every figure draggable; the linear-mappings chapter especially.
- Einsum is All you Need · Tim Rocktäschel · written. Index notation as a power tool, twelve worked examples ending at attention.
- Dive into Deep Learning · Zhang, Lipton, Li & Smola · free book + code. Chapter 2 pairs every operation here with live tensors.
S.4Chapter 03 · From scores to probabilities
- Transformer Explainer · Polo Club · interactive. The temperature/top-k/top-p sliders on a real model; this chapter made tactile.
- ArgMax and SoftMax · StatQuest · video, 14 min. The most concrete softmax arithmetic walk-through on video.
- Temperature, top-k, and top-p · Sebastian Raschka · written, short. The sampling dials with a worked five-token example.
- Seeing Theory · Brown University · interactive. The probability refresher, beautifully built.
- CS231n: Linear Classification · Stanford · course notes. Scores → softmax → loss, the classic bridge into chapter 04.
S.5Chapter 04 · Loss, surprise, and information
S.6Chapter 05 · Gradients: how models learn
S.7Chapter 06 · The transformer block
- Attention, step-by-step · 3Blue1Brown · video, 26 min. The centerpiece recommendation of the whole syllabus.
- The Illustrated Transformer · Jay Alammar · written. The classic diagram walkthrough; the field's shared mental pictures.
- LLM Visualization · Brendan Bycroft · interactive 3-D. Every matrix of the block, drawn to scale, animated mid-inference.
- Rotary Embeddings: A Relative Revolution · EleutherAI · written. The definitive RoPE story, derivation and code.
- The math behind Attention · Luis Serrano · video, 36 min. The gentlest full attention derivation; ideal second angle.
- Let's build GPT · Andrej Karpathy · video, 2 h. The block in working PyTorch, √d experiment included.
- Spreadsheets Are All You Need · Ishan Anand · interactive. GPT-2 in Excel; audit every formula of chapter 06 cell by cell.
- An Interactive Guide to RoPE · Sunil Dhaka · interactive. Draggable rotations; the relative-distance property, felt.
- Rotary Positional Embeddings · Efficient NLP (Bai Li) · video, 11 min. The concise RoPE explainer if you'd rather watch than read.
- Patterns and Messages · Chris McCormick · written series, 2025. The residual-stream-as-bus reframing; the on-ramp to interpretability's vocabulary.
- The Annotated Transformer · Harvard NLP · written + code. The 2017 paper reproduced line by line in PyTorch.
S.8Chapter 07 · The moves papers make
S.9The reference shelf and the graduation reads
- Mathematics for Machine Learning · Deisenroth, Faisal & Ong · free book (Cambridge UP). The rigorous backstop for every chapter here; use as a lookup, not a read-through.
- Dive into Deep Learning · Zhang et al. · free book. Math with runnable tensors, through attention and beyond.
- The Hundred-Page Language Models Book · Andriy Burkov · free book. The best single written companion covering this primer's full arc, foundations to LLMs.
- The Matrix Calculus You Need for Deep Learning · Parr & Howard · written. Jacobians and the vector chain rule, built from calc-1.
- A Visual Guide to Attention Variants in Modern LLMs · Sebastian Raschka, 2026 · written. What attention became: GQA, MLA, sliding windows, hybrids.
- A Primer on the Inner Workings of Transformer-based Language Models · Ferrando et al. · survey. The bridge into interpretability research.
- A Mathematical Framework for Transformer Circuits · Elhage et al., Anthropic · research. The deep end; read it last and recognize everything.
Scoped claim: the gaps this primer covers itself
The research sweep behind this syllabus found no first-rate video teaching for a
few topics, so the primer's own sections and the written or interactive picks above carry them:
perplexity (ch. 04 §4.4, with Huyen's article and TensorTonic's interactive), temperature/top-p
mechanics as one picture (ch. 03 §3.4–3.5, with Raschka's FAQ), and bf16/fp8 number formats
(ch. 07 §7.4; Grootendorst's written guide carries the topic). If you find an excellent video
for any of these, it earns a slot here.