LLM internals · ten chapters · early Transformers → July 2026

Transformers and Memory Systems Evolution

How model memory changed from early Transformers to Kimi K3

Varma Chanderraju · Aug 2026

What is this?

This X article by waterloo_intern on how attention evolved from GPT-2 to Kimi K3 was the trigger for compiling this explainer. Have been meaning to stitch together all the Transformers related references that I have been collecting into a connected piece of prose. Trying to digest the above article prompted me fill in gaps in knowledge about inner workings of Transformers. Used Codex, Claude and Gemini to help me collate everything together. Hope you find it useful. If you find any mistakes or have suggestions for improvement please reach out.

Early decoder-only Transformers relied on exact attention over earlier tokens. Keeping that context became expensive as models grew. Kimi K3 combines several ways to store and retrieve information: a fixed-size recurrent state, a compressed per-token cache, routed expert parameters, and access to earlier layer outputs. Precision and communication costs shape all of them. This explainer follows how those systems developed and compares them with a computer's memory hierarchy. The analogy helps explain why the model uses different stores for different jobs.

The tour ends with Kimi K3context K3's released weights, model card, and technical blog provide enough public detail to examine the systems covered in the explainer. Its inclusion is about available architectural information, not a model ranking., a 2.8-trillion-parameter model released in July 2026. In order to follow along, you need basic matrix multiplication, dot products, and probability distributions. Chapters 01–02 build attention from the beginning. For a refresher on matrices and outer products, see the first three videos of 3Blue1Brown's Essence of Linear Algebra. If you want a deeper dive on the math foundations and have more time to invest review this Math Behind LLMs article.

By the end of this explainer, we should be able to read a model card listing 896 experts, KDA, MLA, Block AttnRes, and MXFP4 and understand what each part does. The explainer has ten chapters, a hands-on lab, a glossary, and live demos throughout.

Plate 0·ATransformers and Memory Systems Evolution
2017–19one exactmemory 2020attention asan RNN 2021delta rulereturns 2023learnedforgetting 2024compressed KV;delta at scale 2025KDA; thehybrid stack 2026depth joins;Kimi K3 …and the six memories it produced: ← small · fast · exact   vast · slow · cheap → SCRATCHPAD recurrent state (KDA) O(1) · fixed · lossy ch. 04–06 LEDGER per-token KV (MLA) compressed · indexed ch. 02–03, 07 LIBRARY expert weights (MoE) 2.8T total · ~4% active ch. 08 DEPTH STREAM residual (AttnRes) retrieval across layers ch. 09 PRECISIONbytes per number: MXFP4/8 quantization-aware training · ch. 10 INTERCONNECTrouting, all-to-all, expert parallelism · ch. 08, 10 Sources: Vaswani 2017 · GPT-2 2019 · Katharopoulos 2020 · Schlag et al. 2021 Mamba 2023 · DeepSeek-V2 and DeltaNet at scale 2024 · Kimi Linear 2025 AttnRes and Kimi K3 2026. Each chapter links to its sources.
The six-part diagram organizes the mechanisms covered here. The memory-and-state framing has earlier sources; each chapter cites the papers behind its mechanisms and dates.

How to read it

Read the chapters in order. Each opens with a memory contract that states what the architecture stores, how that storage grows, what its eviction policy is, and how it fails. The contract changes as the model gains more ways to store and retrieve information:

Memory contract · example (this one belongs to chapter 03)
Stored
One key and one value per past token, per layer, per head, the KV cache
Growth
Linear in context length. Nothing is ever thrown away.
Eviction policy
None; keep everything, reread everything
Failure mode
At long context, memory and bandwidth costs dominate; exactness becomes a luxury

The same colors identify four roles throughout the chapters and figures:

query: "what do I need?" key: "what is stored here?" value: "the content itself" state / write / evict: "the memory being edited"

Dotted-underlined words like KV cache link into the glossary. Dashed-underlined phrases like this onecontext This is a short explanatory note. open short notes labeled context. Hover over or tap them to read more. Longer blue boxes also add context. Amber boxes flag a scoped claim (the limits of a result), and steel boxes explain hardware reality (implementation constraints).

Source materials

The explainer brings existing explanations and primary papers together in one sequence, with interactive demos. Chapters credit the material they use: spotlight cards name sources, quotations appear in marked boxes, adapted figures are labeled, and the embedded videos are the creators' own (3Blue1Brown, Andrej Karpathy). The core references, beyond the primary papers linked where used: 3Blue1Brown, Jay Alammar, Transformer Explainer, Bycroft's LLM Visualization, Karpathy's Zero to Hero, Songlin Yang's DeltaNet series, Raschka's architecture comparison, Grootendorst's MoE guide, and of course the primary source article "22580: From GPT2 to Kimi3, Explained." Claims that exist only in vendor material are labeled vendor-reported, and the interactive demos are toy models that say so.

The chapters

Chapter 01The next-token machineTokens, embeddings, blocks, and the training objective, using GPT-2 small as a concrete example. Chapter 02Attention: exact retrievalHow attention compares a query with earlier tokens and combines their values. Includes an interactive heatmap. Chapter 03The KV cache & the memory wallCalculate the storage and bandwidth cost of exact attention at long context lengths. Chapter 04Linear attention: memory on a budgetStore earlier context in a fixed-size matrix and see how retrieval error grows. Chapter 05DeltaNet: learning to overwriteRead the current value before writing a correction to the state. Chapter 06Forgetting: gates & KDAAdd learned decay to fixed-size memory, then control that decay by channel. Chapter 07The hybrid: KDA + MLACombine fixed-size state with periodic, compressed attention that can address individual tokens. Chapter 08Mixture of ExpertsHow K3 stores 2.8T parameters while activating about 104B per token. Includes a router demo. Chapter 09Attention over depthHow the residual stream accumulates information and AttnRes selects earlier layer outputs. Chapter 10Kimi K3, assembledPut the mechanisms together and check the sources and scope of the published claims. The LabBuild it yourselfEight exercises and Karpathy's build-a-GPT videos. ReferenceGlossaryDefinitions linked from the chapters.

Created by Varma Chanderraju with Codex, Claude, and Gemini.