This explainer starts with the next-token machine and follows the ways language models store, retrieve, compress, and route information as their architectures evolve.

Inside the reader

Foundations 3 chapters
The next-token machine, attention as retrieval, and the KV cache memory wall.
Evolving memory 4 chapters
Linear attention, DeltaNet, gating, Kimi Delta Attention, and hybrid memory designs.
Modern architectures 3 chapters + lab
Mixture of Experts, attention over depth, Kimi K3, a glossary, and hands-on exercises.