Chapter 09 · Retrieval across layers

Attention over depth: AttnRes

Chapters 03–07 dealt with memory across the token sequence. Information also accumulates across layers. In a standard Transformer, each layer adds its output to a running vector. In a deep model, that makes it harder for later layers to use a specific earlier contribution. Attention Residuals (AttnRes) lets layers select earlier outputs instead of relying only on the accumulated sum.

Memory contract · the residual stream (depth memory)
Stored
Every layer's output, summed into one running vector per token, the residual stream
Growth
The vector's size is fixed; what grows with depth is the number of contributions crammed into it
Eviction policy
None; every layer's output is added with weight exactly 1, forever
Failure mode
Dilution: any single layer's contribution shrinks relative to the accumulating whole; later layers can't select what they need

9.1What the residual stream stores

In a PreNormcontext PostNorm normalizes the stream after a layer; PreNorm normalizes each branch's input. PreNorm makes very deep stacks easier to train but leaves the running stream free to grow in magnitude. Transformer, each layer adds its output to the stream: hℓ+1 = hℓ + fℓ(hℓ). Expanding the recurrence shows the sum of all earlier layer contributions:

hL = h0 + Σℓ<L fℓ(hℓ)   · every term with fixed weight 1

Anthropic's interpretability work (A Mathematical Framework for Transformer Circuits, Elhage et al., 2021) taught the field to read this stream as a communication channel between layerscontextIn that picture, heads and MLPs each read from and write into particular directions ("features") of the shared stream. Components must learn how to use that shared representation.. As a memory, the stream has fixed width and receives additive writes from every layer. The Kimi Team's AttnRes paper describes the resulting problem:

Quoted from the "Attention Residuals" abstract · Kimi Team, 2026

Standard residual connections with PreNorm cause "uncontrolled hidden-state growth with depth, progressively diluting each layer's contribution."

arXiv:2603.15031.

Dilution, by the numbersInteractive

Assume (in this simplified modelcontext Real streams can grow in magnitude with depth. The equal-norm assumption leaves that effect out. Lab 6 asks you to measure actual contributions in GPT-2.) every layer contributes a vector of equal norm. The bars show what share of the final stream each layer's contribution represents; uniform addition on the left, a learned softmax weighting (AttnRes-style, peaked on a few useful layers) on the right. Now drag the depth up to K3 scale.

Toy illustration of a real trend: equal-norm contributions are an idealization. The AttnRes paper's actual measurements (norms and gradients across depth) are Figure-level evidence in the paper; this widget only conveys the 1/(L+1) intuition.

In this equal-norm toy, one layer supplies about 8% of the summed contributions at L=12 and about 1% at L=93. Those figures are shares of the sum, not measurements of a real model. If uncorrelated contributions are compared with the stream's norm, which grows roughly as √depth, the relative decline is slower. Either way, a later layer receives the combined vector rather than a direct handle to layer 9's output.

9.2Select earlier layer outputs

Chapter 02 used attention to select information from earlier tokens. AttnRes uses it to select information from earlier layers. The comparison is between layer outputscontext MoE routing in chapter 08 also scores a set of choices, though its top-k selection differs from softmax attention. Here softmax weights the available depth outputs.. AttnRes attends over preceding layer outputs: each layer forms a query and softmax-weights earlier depths, with learned, input-dependent weights replacing the fixed weight-1 addition:

hℓ = Σj<ℓ aℓj vj    Σj aℓj = 1,  aℓj from query–key similarity across depth

Where the pieces come from, concretely: each earlier layer's output is kept as a depth-value with a projected depth-key; the current layer projects a query from its input; and the softmax-weighted sum replaces the plain accumulated stream as what the layer reads. The layer's own computation (attention + MLP) is unchanged; only its input construction is new.

By layer 60, a standard residual stream contains additions from all 59 earlier layers in one vector. Suppose layer 12 produced something useful, but the layers after it added enough that its contribution is hard to pick out. Layer 60 can learn to suppress parts of the mixture, but it has no separate handle for layer 12's output. With AttnRes, earlier outputs, or summaries of blocks of them, remain separate choices. Layer 60 can assign more weight to the earlier contribution it needs. Chapter 02 showed how attention lets a token reach back to a particular earlier token. Here the choice is among earlier layers.

Plate 9·ASame machine, rotated 90°: retrieval across depth instead of time
ch. 02: which TOKENS matter? t−3t−2t−1t attention along the sequence → AttnRes: which DEPTHS matter? h₀ embedding layer 1 output layer 2 output layer 3 output layer ℓ (query) softmax over depths, Σ = 1 ↑ depth attention uses learned weights that change with the input and sum to 1
Mechanism per Attention Residuals (Kimi Team, 2026). The normalized weights also tame the stream's unbounded norm growth: dilution and magnitude growth are two coupled symptoms of the same uniform accumulation.

9.3Block AttnRes limits the storage cost

Keeping every layer's output addressable for every later layer costs activation memory (the stored intermediate states themselves, not parameters) and, in distributed training, where layers live on different devicescontextPipeline parallelism: a 93-layer model is sliced into stages across GPUs, so "attend to layer 9's output" can mean fetching activations from another device. Block-level retrieval reduces the number of such transfers., communication cost. Block AttnRes groups layers and attends over block representations instead of every individual layer output. Per the abstract, the method yields "more uniform output magnitudes and gradient distribution across depth," validated on the Kimi Linear architecture (48B total / 3B activated) trained on 1.4T tokens, with scaling-law experiments across sizescontext Results at several model sizes help test whether an improvement persists as models grow. They do not replace a measurement on every larger model.. Kimi K3 adopts it at fixed block boundaries (Moonshot blog, vendor-reported).

The two kinds of accumulated state can now be compared:

AxisUnmanaged memoryFailureManaged replacement
Time (sequence)Additive linear-attention state (ch. 04)InterferenceDelta rule + gating → KDA (ch. 05–06), rationed exact MLA (ch. 07)
Depth (layers)Additive residual streamDilutionAttnRes; rationed as Block AttnRes

Symmetry framing adapted from the source article by @waterloo_intern.

The memory question applies to depth

The residual stream has a fixed width, accumulates outputs from all earlier layers, and gives each contribution the same initial weight. AttnRes changes how later layers retrieve those contributions. The same questions used for sequence memory—what is stored, how it grows, and what can be selected—help explain this design.

What this chapter established

Go deeper

Contents · Glossary · Quotes from the AttnRes abstract; K3 adoption details are vendor-reported.