Chapter 06 · Controlling decay
DeltaNet can correct an association when it has the right key. To clear broader parts of its fixed-size state, it also needs controlled decay. Gated DeltaNet adds a learned forget gate; Kimi Delta Attention (KDA) lets that gate act differently across channels. Kimi K3 uses KDA in 69 of its 93 attention layerscontext Those 69 layers use linear-time recurrent memory. The other 24 use MLA, as chapter 07 explains..
Chapter 04's state only accumulated writes. Chapter 05 added targeted corrections, but old information elsewhere in the state could still interfere with new information. A decay gate reduces the influence of that older state. The model must learn from the input when to preserve it and when to let it fade.
Imagine the model finishes a legal contract and starts a cooking recipe. The old definitions and clause references may now interfere with new writes. Correcting one key at a time does not clear them as a group. Decay reduces their influence across the state.
That verb comes from the state-space model lineagecontextA parallel research family that reached fixed-state sequence modeling from control theory (continuous dynamical systems, discretized) rather than from attention. In 2024, Mamba's own authors showed SSMs and linear attention are two views of one computation ("Transformers are SSMs," Dao & Gu).. Mamba (Gu & Dao, 2023) made input-dependent decay famous. Mamba's native form is a state-space update, not a key/value outer product, but translated into this explainer's associative-memory notation, the idea it contributes is:
αt ≈ 1: preserve everything. αt ≈ 0: wipe the slate. The gate is input-dependentcontext Earlier recurrent models could use fixed decay schedules. Here the current input helps set the decay strength. The model has to learn which tokens signal that old context matters less.. It can learn to change the gate at topic boundaries.
Two architectures bridged Mamba's gate to the delta rule. RetNet (2023) fixed a decay onto a matrix-valued linear-attention state, and Gated Linear Attention (GLA, Yang et al. 2023) made that decay input-dependent. Gated DeltaNet's move is to pair that gate with chapter 05's targeted overwrite.
Gated Delta Networks (Yang, Kautz & Hatamizadeh, ICLR 2025) put decay and delta correction in one updatecontextWorth noticing that neither verb subsumes the other: decay can't fix a single wrong fact without collateral fading, and a delta write can't clear a thousand stale entries it can't name. Keep / correct / fade: after this chapter, every memory in the explainer speaks this three-verb vocabulary.:
Read it inside-out: decay the whole board (×α), and erase-then-write at the current key: the delta correction acts on the decayed state, using ch. 05's transition form. (Order of α and the erase matrix commutes, since α is a scalar.)
Quoted from the Gated DeltaNet abstract · Yang, Kautz & Hatamizadeh
"gating enables rapid memory erasure while the delta rule facilitates targeted updates"
| Situation in the text | Useful mechanism |
|---|---|
| One stale fact ("the variable was renamed") | Delta correction: β high on the new binding |
| Topic or document boundary | Decay: α drops, whole memory fades |
| Stable long-lived facts | α ≈ 1 and near-zero corrections leave them untouched |
| Contradicted claim | Strong targeted correction at that key |
Table adapted from the source article by @waterloo_intern.
Gated DeltaNet uses one decay value per head and token. KDA lets different channels decay at different rates. A channel is not a stored factcontext Information about one fact can be spread across several channels. Per-channel decay gives the model more control than one rate for the whole state, though it is still coarser than expiring individual facts. Some channels may hold short-lived syntactic information; others may hold document-level information. Kimi Delta Attention, the core operator of Kimi Linear (Kimi Team, 2025), refines the gate to a per-channel decay vector:
The bridge from Gated DeltaNet is one substitution: the scalar α becomes Diag(at), a diagonal matrix holding one learned decay per channel; everything else keeps ch. 05's shape. This is a structurally faithful sketch, not Moonshot's production formulation, which uses further structured transitions and specialized chunkwise kernels (ch. 05's lesson applied); see the paper for the exact form.
Different information can remain useful for different lengths of time. Sentence-level syntax may matter for a few tokens, while the programming language of a code task may matter for hundreds. A user's name should remain available throughout the conversation, though KDA does not guarantee a dedicated slot for it. The model can learn to route components of this information into different channel subspaces, then decay those channels at different rates. The demo shows the effect at a topic change:
The Kimi Linear technical report is a controlled comparison: a 48B-total / 3B-activated hybrid (KDA interleaved with full attention at 3:1, three KDA layers per MLA layer) against a full-MLA baseline under a matched training recipecontext A matched comparison uses the same data, token budget, and tuning effort for both architectures. Otherwise, a performance difference may reflect more than the architecture.. The authors report, in the abstract and official repo:
These results apply to the compared architectures, training recipes, benchmarks, and serving setup. The 6× figure is an "up to" result at 1M context, not a general speed ratio for all KDA models or an end-to-end result for Kimi K3. To compare speedups, check context length, batch size, hardware, and baseline.
What this chapter established
Go deeper
Contents · Glossary · Numbers verified against arXiv abstracts and the MoonshotAI repository, 2026-08-15.