LLM internals · ten chapters · early Transformers → July 2026
How model memory changed from early Transformers to Kimi K3
This X article by waterloo_intern on how attention evolved from GPT-2 to Kimi K3 was the trigger for compiling this explainer. Have been meaning to stitch together all the Transformers related references that I have been collecting into a connected piece of prose. Trying to digest the above article prompted me fill in gaps in knowledge about inner workings of Transformers. Used Codex, Claude and Gemini to help me collate everything together. Hope you find it useful. If you find any mistakes or have suggestions for improvement please reach out.
Early decoder-only Transformers relied on exact attention over earlier tokens. Keeping that context became expensive as models grew. Kimi K3 combines several ways to store and retrieve information: a fixed-size recurrent state, a compressed per-token cache, routed expert parameters, and access to earlier layer outputs. Precision and communication costs shape all of them. This explainer follows how those systems developed and compares them with a computer's memory hierarchy. The analogy helps explain why the model uses different stores for different jobs.
The tour ends with Kimi K3context K3's released weights, model card, and technical blog provide enough public detail to examine the systems covered in the explainer. Its inclusion is about available architectural information, not a model ranking., a 2.8-trillion-parameter model released in July 2026. In order to follow along, you need basic matrix multiplication, dot products, and probability distributions. Chapters 01–02 build attention from the beginning. For a refresher on matrices and outer products, see the first three videos of 3Blue1Brown's Essence of Linear Algebra. If you want a deeper dive on the math foundations and have more time to invest review this Math Behind LLMs article.
By the end of this explainer, we should be able to read a model card listing 896 experts, KDA, MLA, Block AttnRes, and MXFP4 and understand what each part does. The explainer has ten chapters, a hands-on lab, a glossary, and live demos throughout.
Read the chapters in order. Each opens with a memory contract that states what the architecture stores, how that storage grows, what its eviction policy is, and how it fails. The contract changes as the model gains more ways to store and retrieve information:
The same colors identify four roles throughout the chapters and figures:
Dotted-underlined words like KV cache link into the glossary. Dashed-underlined phrases like this onecontext This is a short explanatory note. open short notes labeled context. Hover over or tap them to read more. Longer blue boxes also add context. Amber boxes flag a scoped claim (the limits of a result), and steel boxes explain hardware reality (implementation constraints).
The explainer brings existing explanations and primary papers together in one sequence, with interactive demos. Chapters credit the material they use: spotlight cards name sources, quotations appear in marked boxes, adapted figures are labeled, and the embedded videos are the creators' own (3Blue1Brown, Andrej Karpathy). The core references, beyond the primary papers linked where used: 3Blue1Brown, Jay Alammar, Transformer Explainer, Bycroft's LLM Visualization, Karpathy's Zero to Hero, Songlin Yang's DeltaNet series, Raschka's architecture comparison, Grootendorst's MoE guide, and of course the primary source article "22580: From GPT2 to Kimi3, Explained." Claims that exist only in vendor material are labeled vendor-reported, and the interactive demos are toy models that say so.
Created by Varma Chanderraju with Codex, Claude, and Gemini.