Chapter 07 · Combining two kinds of memory

The hybrid: KDA + MLA

Chapter 06 gave KDA control over what to keep, correct, and forget. That control cannot make a small, fixed-size state hold every detail of a million-token context. For scale, a 64×64 fp16 matrix holds 8 KBcontext A million tokens contain far more information than that matrix can hold. Compression must discard some detail, even when the model learns which details to keep.. If an answer depends on the exact token from 800,000 positions earlier, that summary may have lost it. Kimi K3 therefore combines recurrent KDA with periodic MLA layers that keep a compressed but individually addressable slot for each token.

Memory contract · hybrid KDA + MLA
Stored
Most layers: fixed-size KDA state. Every 4th layer: a token-indexed KV cache, compressed into small latent vectors (MLA)
Growth
Constant in the KDA layers; linear (but ~4× fewer layers and far fewer bytes per token) in the MLA layers
Eviction policy
KDA layers: decay + overwrite. MLA layers: none, and that's the point; they are the ledger of record, every token keeping its own (compressed) slot
Failure mode
Full-context, token-addressed attention is now rationed: available, but only at periodic layers

7.1When approximate recall fails

These tasks need precise retrieval:

Task list adapted from the source article by @waterloo_intern.

A fixed state can retain broad information, but it has no separate address for each earlier token. You cannot ask it to reread position 214,067 exactly. These tasks need addressability. In the recall literature this shows up as associative-recall and needle-in-a-haystackcontextThe benchmark is what it sounds like: hide one specific fact ("the passcode is 7482") deep inside hundreds of thousands of tokens of filler, then ask for it. Treat single-needle scores with mild suspicion: models can pass while still failing multi-fact or reasoning-over-context variants. But as a lower bound it's clarifying: a memory that can't do this can't do verbatim recall, period. benchmarks, where pure linear-attention models have historically struggled (as discussed in Songlin Yang's series). The hybrid keeps some layers with an address for every token.

Think of a document summary and an indexed copy of its pages. The summary can help answer "what is this contract about?" To reproduce clause 14(b) exactly, you need to find the clause itself. These are different abilities: one keeps a compact, lossy view; the other keeps an address for each part of the source. Improving the summary cannot give it addresses for passages it no longer holds. A model expected to handle both kinds of long-context question needs both abilities. Kimi combines them without keeping a full KV cache at every layer; MLA preserves the addresses while compressing what each slot stores.

7.2MLA: token-addressable retrieval, compressed

The per-token cache can still be compressed. That is Multi-head Latent Attention, introduced by DeepSeek-V2 (2024). Instead of caching full keys and values for every head, MLA learns a joint low-rank compressioncontext The full K/V tensors across heads contain redundant information. MLA learns, during training, to encode them in a smaller shared vector.: each token's K/V information is squeezed through a small latent vector, and per-head keys and values are re-expanded from it when neededcontextOne real wrinkle: rotary position encoding doesn't survive the squeeze: rotation and low-rank projection don't commute, so MLA carries a small separate positional key alongside the latent (the "decoupled RoPE" of the DeepSeek-V2 paper, §2). The kind of asterisk that separates paper architectures from shipping ones. One more: on the decode path the up-projections fold into the query and output weights, so the per-head keys and values are never materialized in the cache, which holds only the latent..

Addressing and stored content change in different ways. Softmax attention still has one slot per token, so it can single out position 214,067. Each slot holds a compressed representation, which can lose information. DeepSeek-V2's results suggest that much of the original K/V representation was redundant.

Plate 7·AWhat sits in the cache: MHA vs. GQA vs. MLA
MHA · per token: one K + V per head full cache cost (ch. 03) GQA · per token: few K/V heads shared by many query heads. Same idea, coarser knife. MLA · per token: one small latent vector, jointly encoding K and V for ALL heads ← re-expanded on the fly by learned up-projections DeepSeek-V2 reports this cuts its KV cache by 93.3% vs. its 67B (GQA) predecessor, while keeping every token individually addressable. Compression ≠ forgetting: each token keeps a slot of its own.
Mechanism and the 93.3% figure from the DeepSeek-V2 paper (abstract), verified 2026-08-15. For a code-first build-it-yourself treatment, see PyImageSearch's MLA tutorial or Jinpeng Zhang's technical analysis.

MLA compresses each slotcontext GQA shares K/V projections across heads, sliding-window attention drops older slots, MLA makes each slot smaller, and linear attention replaces the slots with fixed state. K3 combines MLA and KDA in different layers.. Its cache still grows linearly with context length, but uses fewer bytes per token. KDA keeps a fixed-size state. Their strengths and limits are different:

ComponentMemory formStrengthWeakness
KDAFixed-size recurrent matrixO(1) state; cheap million-token decodeLossy: gist, not verbatim
MLAToken-indexed compressed KV cachePer-token, addressable retrievalCache and read cost still grow with T

Framing adapted from the source article.

7.3How the two layers alternate

Kimi Linear repeats three KDA layers followed by one MLA layercontext The 3:1 ratio is the design reported by the Kimi team. K3 has 69 KDA and 24 MLA layers, or about 2.9:1.. The official repository states this ratio. The KDA layers use fixed-size state; the periodic MLA layer can look up a specific earlier token across the full context.

Plate 7·BOne macro-cycle of the hybrid stack
KDA · fixed-state sequence memory KDA KDA MLA · full-context, token-addressed attention (the only layer keeping a cache) …then the cycle repeats up the stack tokens flow upward
Ratio per the MoonshotAI/Kimi-Linear repository. Kimi K3 keeps the same shape at scale: 69 KDA + 24 gated-MLA layers (model card), ch. 10.

The layer ratio helps explain the reported up-to-75% reduction in KV-cache use: three quarters of the layers keep no per-token cache, and the MLA layers compress theirs. The reported up-to-6× decode throughput at 1M context is a measured serving result. It also depends on kernels, bandwidth, batching, and the comparison setup.

Why the layers are mixed

Exact attention preserves token-level lookup but costs more memory. Using it once per four-layer cycle keeps that ability while most layers use fixed-size state. Chapter 08 applies selective use to expert parameters, and chapter 09 examines selective access to earlier layers.

What this chapter established

Go deeper

Contents · Glossary · MLA figures from the DeepSeek-V2 abstract; ratios from MoonshotAI sources, verified 2026-08-15.