Chapter 10 · Kimi K3's memory systems
Kimi K3 combines the mechanisms from the previous chapters. The components address different limitscontext KDA uses fixed-size but lossy state; periodic MLA layers preserve token-level lookup. MoE stores far more parameters than it activates for one token. AttnRes lets deep layers select earlier outputs. Together they balance retrieval accuracy, arithmetic work, memory bandwidth, capacity, and communication across devices. This chapter puts the published specifications in one place and separates reported measurements from architectural facts.
Each row names its source. The model card and technical blog are Moonshot publications and describe the model's design. Their large-scale efficiency and quality claims have not all been independently reproducedcontext Checking a model of this size can mean running released weights, reproducing smaller experiments, or measuring serving performance. Retraining all 2.8T parameters independently is much more costly..
| Quantity | Value | Source |
|---|---|---|
| Total parameters | 2.8T | Model card |
| Activated per token | 104B | Model card |
| Experts | 896 routed · 16 selected/token · 2 shared | Model card |
| Attention stack | 69 KDA + 24 gated MLA = 93 layers (≈ 3:1, the Kimi Linear recipe at scale) | Model card; ratio design per Kimi-Linear repo |
| Context window | 1,048,576 tokens | Model card |
| Attention hidden dim / heads | 7,168 / 96 | Model card |
| Dense layers | 1 (all other FFN positions are MoE) | Model card |
| Vocabulary | 160K | Model card |
| Vision encoder | MoonViT-V2, 401M (native multimodal) | Model card |
| Precision | MXFP4 weights / MXFP8 activations, QAT from the SFT stage onward | Model card + blog |
| Expert framework / activation | Stable LatentMoE · SiTU-GLU | Blog + model card (vendor-reported design detail) |
All values re-checked against the linked sources on 2026-08-15.
The six systems serve different purposes (adapted from the source article):
K3 trains quantization-aware from the SFT stagecontextSupervised fine-tuning, the post-pretraining stage where the model learns from curated instruction/response examples. Starting QAT here is a pragmatic compromise: pretraining still runs in higher precision (stability). The later assistant-training stage exposes the model to rounding used by the serving format.: the model learns while experiencing the rounding of its serving format, MXFP4context"MX" is microscaling. Small blocks of numbers share a scale factor, giving their 4-bit floats a wider useful range than each could represent alone. for weights and MXFP8 for activations. Bytes per number also determine memory use. A 2.8T-parameter model at 4 bits per weight needs roughly a quarter of the raw weight storage of fp16 (before block-scale metadata and any tensors kept at higher precision). Details are vendor-reported in the blog and model card.
The source article closes with a useful checklist. For claims such as "6× faster" or "beats full attention," check:
Checklist items consolidated from the caveats sections of the source article.
The six systems resemble parts of a computer's memory hierarchy. The analogy has an important limitcontextThe deepest mismatch: a CPU hierarchy caches the same bytes at different distances, while the LLM systems hold different representations: a compressed sequence state, token-addressed details, and learned expert parameters. They are not copies of one underlying set of bytes. The table compares their roles:
| Classic hierarchy | K3's version | The twist |
|---|---|---|
| Registers | KDA recurrent state: tiny, touched every step, fixed size | Eviction is learned (α, β per token), not hardwired |
| Cache / RAM | MLA's compressed KV cache: addressable, grows with the working set | The compression codec is trained with the model |
| Paged storage | MoE expert banks: vast, resident across devices; tokens travel to their selected experts | The "page table" is a neural router choosing 16 of 896 (though nothing is actually paged in) |
| Forwarding / bypass paths | Block AttnRes across depth | Which earlier stage to forward from is itself attention |
| Word size | MXFP4 / MXFP8 microscaling formats | Chosen during training, not after |
| The bus | Expert-parallel all-to-all interconnect | Topology constraints shape the architecture, not just its speed |
Classical hierarchies often use fixed management policies. Several K3 choices, including KDA's decay and MoE's routing, depend on the current input and were learned during training. The table is an analogy for those choices, not a claim that K3 literally pages expert weights like a CPU pages memory.
Comparing LLM memory with a computer's hierarchy is not new: PagedAttention (ch. 03) manages the KV cache the way an operating system pages virtual memory, and MemGPT applies an OS memory hierarchy to context management. This table maps that on to K3's subsystems and notes the limits of that comparison.
| Ch. | Mechanism | Problem it solved | New limit it exposed |
|---|---|---|---|
| 02 | Softmax attention | Exact token-to-token retrieval | Cost grows with context |
| 03 | KV cache | No recomputation during decode | Linear memory, bandwidth-bound decode |
| 04 | Linear attention | Fixed-size recurrent memory | Interference in finite state |
| 05 | Delta rule | Targeted replacement of associations | No cheap wholesale forgetting |
| 06 | Gating → KDA | Broad decay; then per-channel decay | Still lossy: finite state can't be verbatim |
| 07 | Hybrid + MLA | Rationed exact retrieval, compressed slots | Exactness is now a budget line |
| 08 | Sparse MoE | Capacity without per-token compute | Routing, communication, balance |
| 09 | Block AttnRes | Selective retrieval across depth | Extra state and communication |
Table structure from the source article's "tutorial in miniature," extended with chapter mapping.
Will learned finite-state memory with periodic exact retrieval become common in frontier models, or will better sparse and full-attention kernels keep exact lookup affordable across more layers? Kimi Linear and Gated DeltaNet offer controlled evidence for the first approach, within the models and setups they tested. K3 is a much larger implementation of that design direction.
Each system reduces a different cost. KDA and MLA limit how much token-addressable context the model keeps. MoE limits active parameters per token. MXFP4 reduces weight storage, Block AttnRes limits access across depth, and expert parallelism has to manage communication between devices. These are distinct resource constraints.
K3 reflects the costs of current hardware: memory bandwidth, arithmetic relative to data movement, and transfers between accelerators. If exact attention becomes cheaper, the case for KDA and MLA could change. More abundant memory and interconnect bandwidth could change the value of aggressive quantization and expert routing. Parameter count and depth-related training issues may remain. When you read a new model card, check which costs its design reduces and which costs it accepts.
What the explainer established
Go deeper: the primary sources, in reading order
Contents · Glossary · Spec values verified against the model card and blog on 2026-08-15; vendor-reported items labeled throughout.