This explainer starts with the next-token machine and follows the ways language models store, retrieve, compress, and route information as their architectures evolve.
Inside the reader
- Foundations 3 chapters
- The next-token machine, attention as retrieval, and the KV cache memory wall.
- Evolving memory 4 chapters
- Linear attention, DeltaNet, gating, Kimi Delta Attention, and hybrid memory designs.
- Modern architectures 3 chapters + lab
- Mixture of Experts, attention over depth, Kimi K3, a glossary, and hands-on exercises.