Chapter 10 · Kimi K3's memory systems

Kimi K3, assembled

Kimi K3 combines the mechanisms from the previous chapters. The components address different limitscontext KDA uses fixed-size but lossy state; periodic MLA layers preserve token-level lookup. MoE stores far more parameters than it activates for one token. AttnRes lets deep layers select earlier outputs. Together they balance retrieval accuracy, arithmetic work, memory bandwidth, capacity, and communication across devices. This chapter puts the published specifications in one place and separates reported measurements from architectural facts.

Memory contract · Kimi K3
Sequence memory
69 KDA layers (fixed state, decay + delta eviction) + 24 gated-MLA layers (compressed, token-addressable retrieval), ch. 04–07
Parameter memory
2.8T stored; ~104B active per token; Stable LatentMoE routing is the main sparsity mechanism, ch. 08
Depth memory
Block AttnRes: learned retrieval over layer blocks, ch. 09
Precision
MXFP4 weights / MXFP8 activations, quantization-aware from SFT onward

10.1The spec sheet, with provenance

Each row names its source. The model card and technical blog are Moonshot publications and describe the model's design. Their large-scale efficiency and quality claims have not all been independently reproducedcontext Checking a model of this size can mean running released weights, reproducing smaller experiments, or measuring serving performance. Retraining all 2.8T parameters independently is much more costly..

QuantityValueSource
Total parameters2.8TModel card
Activated per token104BModel card
Experts896 routed · 16 selected/token · 2 sharedModel card
Attention stack69 KDA + 24 gated MLA = 93 layers (≈ 3:1, the Kimi Linear recipe at scale)Model card; ratio design per Kimi-Linear repo
Context window1,048,576 tokensModel card
Attention hidden dim / heads7,168 / 96Model card
Dense layers1 (all other FFN positions are MoE)Model card
Vocabulary160KModel card
Vision encoderMoonViT-V2, 401M (native multimodal)Model card
PrecisionMXFP4 weights / MXFP8 activations, QAT from the SFT stage onwardModel card + blog
Expert framework / activationStable LatentMoE · SiTU-GLUBlog + model card (vendor-reported design detail)

All values re-checked against the linked sources on 2026-08-15.

10.2One macro-cycle, all six systems

Plate 10·AThe K3 macro-cycle: every chapter, in one column
tokens (and image patches) flow upward embeddings · text tokens + MoonViT-V2 vision features KDA · fixed-state sequence memory (ch. 04–06) Stable LatentMoE · 16/896 experts + 2 shared (ch. 08) KDA Stable LatentMoE KDA Stable LatentMoE Gated MLA · periodic token-addressed attention (ch. 07) Block AttnRes · selective read across depth (ch. 09) …then the cycle repeats, 93 attention layers in total ×69 ×24
Layer schematic adapted from the source article's architecture walkthrough, corrected against the model card layer counts. Exact block boundaries within the stack are simplified, and Block AttnRes is drawn as what it is: a change to how blocks read the accumulated stream (dashed side paths), not an extra layer in the stack. See Moonshot's blog for their diagrams.

The six systems serve different purposes (adapted from the source article):

  1. KDA: cheap, lossy, managed recurrent sequence memory;
  2. Gated MLA: expensive, compressed, token-addressable retrieval, rationed 1-in-4;
  3. Stable LatentMoE: conditional parameter capacity, 16-of-896;
  4. Block AttnRes: selective depth retrieval;
  5. Quantization-aware training: a low-precision serving path built in, not bolted on;
  6. Expert-parallel infrastructure: the routing, balancing, and communication engineering that makes 1–5 executable.

10.3Precision changes memory use

K3 trains quantization-aware from the SFT stagecontextSupervised fine-tuning, the post-pretraining stage where the model learns from curated instruction/response examples. Starting QAT here is a pragmatic compromise: pretraining still runs in higher precision (stability). The later assistant-training stage exposes the model to rounding used by the serving format.: the model learns while experiencing the rounding of its serving format, MXFP4context"MX" is microscaling. Small blocks of numbers share a scale factor, giving their 4-bit floats a wider useful range than each could represent alone. for weights and MXFP8 for activations. Bytes per number also determine memory use. A 2.8T-parameter model at 4 bits per weight needs roughly a quarter of the raw weight storage of fp16 (before block-scale metadata and any tensors kept at higher precision). Details are vendor-reported in the blog and model card.

10.4Questions to ask about performance claims

The source article closes with a useful checklist. For claims such as "6× faster" or "beats full attention," check:

Checklist items consolidated from the caveats sections of the source article.

10.5Compare the systems with a memory hierarchy

The six systems resemble parts of a computer's memory hierarchy. The analogy has an important limitcontextThe deepest mismatch: a CPU hierarchy caches the same bytes at different distances, while the LLM systems hold different representations: a compressed sequence state, token-addressed details, and learned expert parameters. They are not copies of one underlying set of bytes. The table compares their roles:

Classic hierarchyK3's versionThe twist
RegistersKDA recurrent state: tiny, touched every step, fixed sizeEviction is learned (α, β per token), not hardwired
Cache / RAMMLA's compressed KV cache: addressable, grows with the working setThe compression codec is trained with the model
Paged storageMoE expert banks: vast, resident across devices; tokens travel to their selected expertsThe "page table" is a neural router choosing 16 of 896 (though nothing is actually paged in)
Forwarding / bypass pathsBlock AttnRes across depthWhich earlier stage to forward from is itself attention
Word sizeMXFP4 / MXFP8 microscaling formatsChosen during training, not after
The busExpert-parallel all-to-all interconnectTopology constraints shape the architecture, not just its speed

Classical hierarchies often use fixed management policies. Several K3 choices, including KDA's decay and MoE's routing, depend on the current input and were learned during training. The table is an analogy for those choices, not a claim that K3 literally pages expert weights like a CPU pages memory.

Comparing LLM memory with a computer's hierarchy is not new: PagedAttention (ch. 03) manages the KV cache the way an operating system pages virtual memory, and MemGPT applies an OS memory hierarchy to context management. This table maps that on to K3's subsystems and notes the limits of that comparison.

10.6The whole road, one table

Ch.MechanismProblem it solvedNew limit it exposed
02Softmax attentionExact token-to-token retrievalCost grows with context
03KV cacheNo recomputation during decodeLinear memory, bandwidth-bound decode
04Linear attentionFixed-size recurrent memoryInterference in finite state
05Delta ruleTargeted replacement of associationsNo cheap wholesale forgetting
06Gating → KDABroad decay; then per-channel decayStill lossy: finite state can't be verbatim
07Hybrid + MLARationed exact retrieval, compressed slotsExactness is now a budget line
08Sparse MoECapacity without per-token computeRouting, communication, balance
09Block AttnResSelective retrieval across depthExtra state and communication

Table structure from the source article's "tutorial in miniature," extended with chapter mapping.

An open architecture question

Will learned finite-state memory with periodic exact retrieval become common in frontier models, or will better sparse and full-attention kernels keep exact lookup affordable across more layers? Kimi Linear and Gated DeltaNet offer controlled evidence for the first approach, within the models and setups they tested. K3 is a much larger implementation of that design direction.

10.7Why this design may change

Each system reduces a different cost. KDA and MLA limit how much token-addressable context the model keeps. MoE limits active parameters per token. MXFP4 reduces weight storage, Block AttnRes limits access across depth, and expert parallelism has to manage communication between devices. These are distinct resource constraints.

K3 reflects the costs of current hardware: memory bandwidth, arithmetic relative to data movement, and transfers between accelerators. If exact attention becomes cheaper, the case for KDA and MLA could change. More abundant memory and interconnect bandwidth could change the value of aggressive quantization and expert routing. Parameter count and depth-related training issues may remain. When you read a new model card, check which costs its design reduces and which costs it accepts.

What the explainer established

Go deeper: the primary sources, in reading order

Contents · Glossary · Spec values verified against the model card and blog on 2026-08-15; vendor-reported items labeled throughout.