Chapter 09 · Retrieval across layers
Chapters 03–07 dealt with memory across the token sequence. Information also accumulates across layers. In a standard Transformer, each layer adds its output to a running vector. In a deep model, that makes it harder for later layers to use a specific earlier contribution. Attention Residuals (AttnRes) lets layers select earlier outputs instead of relying only on the accumulated sum.
In a PreNormcontext PostNorm normalizes the stream after a layer; PreNorm normalizes each branch's input. PreNorm makes very deep stacks easier to train but leaves the running stream free to grow in magnitude. Transformer, each layer adds its output to the stream: hℓ+1 = hℓ + fℓ(hℓ). Expanding the recurrence shows the sum of all earlier layer contributions:
Anthropic's interpretability work (A Mathematical Framework for Transformer Circuits, Elhage et al., 2021) taught the field to read this stream as a communication channel between layerscontextIn that picture, heads and MLPs each read from and write into particular directions ("features") of the shared stream. Components must learn how to use that shared representation.. As a memory, the stream has fixed width and receives additive writes from every layer. The Kimi Team's AttnRes paper describes the resulting problem:
Quoted from the "Attention Residuals" abstract · Kimi Team, 2026
Standard residual connections with PreNorm cause "uncontrolled hidden-state growth with depth, progressively diluting each layer's contribution."
In this equal-norm toy, one layer supplies about 8% of the summed contributions at L=12 and about 1% at L=93. Those figures are shares of the sum, not measurements of a real model. If uncorrelated contributions are compared with the stream's norm, which grows roughly as √depth, the relative decline is slower. Either way, a later layer receives the combined vector rather than a direct handle to layer 9's output.
Chapter 02 used attention to select information from earlier tokens. AttnRes uses it to select information from earlier layers. The comparison is between layer outputscontext MoE routing in chapter 08 also scores a set of choices, though its top-k selection differs from softmax attention. Here softmax weights the available depth outputs.. AttnRes attends over preceding layer outputs: each layer forms a query and softmax-weights earlier depths, with learned, input-dependent weights replacing the fixed weight-1 addition:
Where the pieces come from, concretely: each earlier layer's output is kept as a depth-value with a projected depth-key; the current layer projects a query from its input; and the softmax-weighted sum replaces the plain accumulated stream as what the layer reads. The layer's own computation (attention + MLP) is unchanged; only its input construction is new.
By layer 60, a standard residual stream contains additions from all 59 earlier layers in one vector. Suppose layer 12 produced something useful, but the layers after it added enough that its contribution is hard to pick out. Layer 60 can learn to suppress parts of the mixture, but it has no separate handle for layer 12's output. With AttnRes, earlier outputs, or summaries of blocks of them, remain separate choices. Layer 60 can assign more weight to the earlier contribution it needs. Chapter 02 showed how attention lets a token reach back to a particular earlier token. Here the choice is among earlier layers.
Keeping every layer's output addressable for every later layer costs activation memory (the stored intermediate states themselves, not parameters) and, in distributed training, where layers live on different devicescontextPipeline parallelism: a 93-layer model is sliced into stages across GPUs, so "attend to layer 9's output" can mean fetching activations from another device. Block-level retrieval reduces the number of such transfers., communication cost. Block AttnRes groups layers and attends over block representations instead of every individual layer output. Per the abstract, the method yields "more uniform output magnitudes and gradient distribution across depth," validated on the Kimi Linear architecture (48B total / 3B activated) trained on 1.4T tokens, with scaling-law experiments across sizescontext Results at several model sizes help test whether an improvement persists as models grow. They do not replace a measurement on every larger model.. Kimi K3 adopts it at fixed block boundaries (Moonshot blog, vendor-reported).
The two kinds of accumulated state can now be compared:
| Axis | Unmanaged memory | Failure | Managed replacement |
|---|---|---|---|
| Time (sequence) | Additive linear-attention state (ch. 04) | Interference | Delta rule + gating → KDA (ch. 05–06), rationed exact MLA (ch. 07) |
| Depth (layers) | Additive residual stream | Dilution | AttnRes; rationed as Block AttnRes |
Symmetry framing adapted from the source article by @waterloo_intern.
The residual stream has a fixed width, accumulates outputs from all earlier layers, and gives each contribution the same initial weight. AttnRes changes how later layers retrieve those contributions. The same questions used for sequence memory—what is stored, how it grows, and what can be selected—help explain this design.
What this chapter established
Go deeper
Contents · Glossary · Quotes from the AttnRes abstract; K3 adoption details are vendor-reported.