This visual primer builds the mathematical intuition needed to follow how language models learn and how a transformer processes information. It is meant as a practical refresher rather than a formal math course.

Inside the reader

Foundations 5 chapters
Vectors, matrices, probability, loss, gradients, and the role each plays in model training.
Transformers 1 chapter
A guided pass through attention, MLPs, residual connections, and normalization.
Reading papers 1 chapter + references
Common mathematical moves, a Greek-letter guide, glossary, and curated learning resources.