Reference · every term, defined once

Glossary

Definitions are alphabetical. Each entry points to the chapter that explains the concept in detail.

all-to-all communication ch. 08
The exchange in expert-parallel MoE where devices send tokens to the devices holding their selected experts and receive the results. Communication can limit the benefit of sparse computation.
AttnRes (Attention Residuals) ch. 09
Replacing the residual stream's fixed weight-1 accumulation with softmax attention over preceding layer outputs: learned, input-dependent retrieval across depth. Kimi Team, arXiv:2603.15031.
bandwidth-bound ch. 03
A computation limited by how fast data moves from memory, not by arithmetic throughput. Cached LLM decode is the canonical example: little math per byte reread.
Block AttnRes ch. 09
A version of AttnRes that groups layers into blocks and attends over block representations. It uses less memory and communication than addressing every layer separately.
causal mask ch. 02
The −∞ pattern added to attention scores so each token can only attend to earlier positions. What makes parallel training compatible with left-to-right generation.
channel ch. 06
One dimension of a hidden vector or state matrix. KDA can set a different decay rate for each channel.
chunkwise parallelism ch. 05
Computing a recurrence in chunks: dense parallel math within each chunk, sequential state hand-off between chunks. More FLOPs than the pure recurrence, far faster on GPUs.
decay gate (α) ch. 06
A learned multiplier ≤ 1 that reduces earlier state at each step. Gated DeltaNet uses one value per head; KDA uses a value per channel.
decode ch. 03
The one-token-at-a-time generation phase, reading the KV cache each step. Bandwidth-bound, in contrast to prefill.
delta rule ch. 05
Error-correction update from Widrow & Hoff (1960): read what the memory returns for a key, write only β × (target − retrieved). Revision instead of accumulation.
dilution ch. 09
The shrinking relative weight of any single layer's contribution in a deep residual stream, since all contributions are summed with weight 1.
embedding ch. 01
The learned mapping from a token ID to a vector. GPT-2 also uses the embedding table at its output head.
expert / shared expert ch. 08
One of many parallel MLPs in an MoE layer. Routed experts serve only tokens sent to them; shared experts process every token, holding common knowledge so routed ones can specialize.
fast weight programmer ch. 05
1990s Schmidhuber concept, revived 2021: a slow (trained) network emits keys/values/rates that program a fast, temporary weight matrix during the forward pass. Linear attention is formally one.
feature map φ ch. 04
The function applied separately to queries and keys in linear attention, replacing softmax so the computation can be re-associated into a recurrent state.
FLOPs ch. 03, 05
Floating-point operations. FLOP count measures arithmetic work, which does not by itself determine run time.
GQA (grouped-query attention) ch. 03, 07
Several query heads share each K/V head, reducing KV-cache size in proportion to the number of shared heads.
interference / overcapacity ch. 04
Cross-talk between associations stored in a fixed-size additive memory once their count approaches the state dimension: only d vectors can be mutually orthogonal in d dimensions.
KDA (Kimi Delta Attention) ch. 06
Gated DeltaNet with fine-grained, channel-wise decay and a hardware-efficient chunkwise formulation. Runs 69 of Kimi K3's 93 attention layers. Kimi Linear, arXiv:2510.26692.
kernel (GPU) ch. 05, 08
A compiled function launched on a GPU. Fusing adjacent operations into one kernel can reduce memory transfers and launch overhead.
KV cache ch. 03
The stored keys and values of all past tokens, kept so decode doesn't recompute them. Grows linearly with context; the central cost of long-context exact attention.
LatentMoE (Stable LatentMoE) ch. 08
K3's MoE framework (vendor-reported): token representations are compressed to a latent space before expert computation and re-expanded after, cutting expert FLOPs and communication.
logits ch. 01
The raw per-vocabulary-token scores produced by the output head, turned into probabilities by softmax.
MLA (multi-head latent attention) ch. 07
Attention that stores a small, jointly compressed K/V vector for each token. The model can still address individual tokens, though each slot holds compressed content. In optimized decoding, up-projections fold into the query and output weights, so full per-head K/V tensors need not be stored. DeepSeek-V2, arXiv:2405.04434.
MoE (Mixture of Experts) ch. 08
Replacing each dense MLP with many experts plus a router that activates a few per token, multiplying stored capacity while holding per-token compute nearly constant.
outer product ch. 04
kv⊤: a d×d rank-1 matrix that associates key direction k with value v. Linear attention adds these matrices to its state.
prefill ch. 03
Processing the whole prompt in parallel before generation starts. Compute-bound and quadratic in prompt length.
PreNorm ch. 09
Placing layer normalization before each sub-layer, leaving the residual stream itself un-normalized, the standard recipe, and the reason stream magnitude grows with depth.
quantization-aware training (QAT) ch. 10
Training (or fine-tuning) while simulating the low-precision number format used at serving time, so the model adapts to the rounding. K3: MXFP4 weights / MXFP8 activations from SFT onward.
query / key / value ch. 02
The three learned projections of attention: what am I looking for / what is stored here / the content itself. Color-coded this way throughout the explainer.
residual stream ch. 01, 09
The running vector for each token to which layers add their outputs. It carries accumulated information across depth; the term comes from Anthropic's interpretability work.
RNN / recurrent state ch. 04
A model that carries a fixed-size state forward step by step. Linear attention is attention re-derived as an RNN whose state is a matrix.
router / load balancing / capacity factor ch. 08
The small scoring network choosing top-k experts per token; the auxiliary pressure keeping traffic spread across experts; and the per-expert buffer size beyond which tokens overflow.
SiTU ch. 08
"Sigmoid Tanh Unit": the gated activation in K3's expert path (SiTU-GLU in the model card). Vendor-reported design detail; its lesson here is about kernel fusion, not the formula.
softmax ch. 01, 02
Exponentiate and normalize: turns a score vector into a probability distribution. Sharpens gaps between scores; see ch. 02's sharpness slider.
state-space model / Mamba ch. 06
The parallel lineage of fixed-state sequence models from which input-dependent decay gates entered the story. Gu & Dao, arXiv:2312.00752.
token ch. 01
A vocabulary chunk (word piece), the unit models read, predict, and measure context length in.

Contents · Definitions compressed from the chapters, which credit their sources inline.