Reference · every term, defined once
Glossary
Definitions are alphabetical. Each entry points to the chapter
that explains the concept in detail.
- all-to-all communication ch. 08
- The exchange in expert-parallel MoE where devices send tokens to the devices holding their selected experts and receive the results. Communication can limit the benefit of sparse computation.
- AttnRes (Attention Residuals) ch. 09
- Replacing the residual stream's fixed weight-1 accumulation with softmax attention over preceding layer outputs: learned, input-dependent retrieval across depth. Kimi Team, arXiv:2603.15031.
- bandwidth-bound ch. 03
- A computation limited by how fast data moves from memory, not by arithmetic throughput. Cached LLM decode is the canonical example: little math per byte reread.
- Block AttnRes ch. 09
- A version of AttnRes that groups layers into blocks and attends over block representations. It uses less memory and communication than addressing every layer separately.
- causal mask ch. 02
- The −∞ pattern added to attention scores so each token can only attend to earlier positions. What makes parallel training compatible with left-to-right generation.
- channel ch. 06
- One dimension of a hidden vector or state matrix. KDA can set a different decay rate for each channel.
- chunkwise parallelism ch. 05
- Computing a recurrence in chunks: dense parallel math within each chunk, sequential state hand-off between chunks. More FLOPs than the pure recurrence, far faster on GPUs.
- decay gate (α) ch. 06
- A learned multiplier ≤ 1 that reduces earlier state at each step. Gated DeltaNet uses one value per head; KDA uses a value per channel.
- decode ch. 03
- The one-token-at-a-time generation phase, reading the KV cache each step. Bandwidth-bound, in contrast to prefill.
- delta rule ch. 05
- Error-correction update from Widrow & Hoff (1960): read what the memory returns for a key, write only β × (target − retrieved). Revision instead of accumulation.
- dilution ch. 09
- The shrinking relative weight of any single layer's contribution in a deep residual stream, since all contributions are summed with weight 1.
- embedding ch. 01
- The learned mapping from a token ID to a vector. GPT-2 also uses the embedding table at its output head.
- expert / shared expert ch. 08
- One of many parallel MLPs in an MoE layer. Routed experts serve only tokens sent to them; shared experts process every token, holding common knowledge so routed ones can specialize.
- fast weight programmer ch. 05
- 1990s Schmidhuber concept, revived 2021: a slow (trained) network emits keys/values/rates that program a fast, temporary weight matrix during the forward pass. Linear attention is formally one.
- feature map φ ch. 04
- The function applied separately to queries and keys in linear attention, replacing softmax so the computation can be re-associated into a recurrent state.
- FLOPs ch. 03, 05
- Floating-point operations. FLOP count measures arithmetic work, which does not by itself determine run time.
- GQA (grouped-query attention) ch. 03, 07
- Several query heads share each K/V head, reducing KV-cache size in proportion to the number of shared heads.
- interference / overcapacity ch. 04
- Cross-talk between associations stored in a fixed-size additive memory once their count approaches the state dimension: only d vectors can be mutually orthogonal in d dimensions.
- KDA (Kimi Delta Attention) ch. 06
- Gated DeltaNet with fine-grained, channel-wise decay and a hardware-efficient chunkwise formulation. Runs 69 of Kimi K3's 93 attention layers. Kimi Linear, arXiv:2510.26692.
- kernel (GPU) ch. 05, 08
- A compiled function launched on a GPU. Fusing adjacent operations into one kernel can reduce memory transfers and launch overhead.
- KV cache ch. 03
- The stored keys and values of all past tokens, kept so decode doesn't recompute them. Grows linearly with context; the central cost of long-context exact attention.
- LatentMoE (Stable LatentMoE) ch. 08
- K3's MoE framework (vendor-reported): token representations are compressed to a latent space before expert computation and re-expanded after, cutting expert FLOPs and communication.
- logits ch. 01
- The raw per-vocabulary-token scores produced by the output head, turned into probabilities by softmax.
- MLA (multi-head latent attention) ch. 07
- Attention that stores a small, jointly compressed K/V vector for each token. The model can still address individual tokens, though each slot holds compressed content. In optimized decoding, up-projections fold into the query and output weights, so full per-head K/V tensors need not be stored. DeepSeek-V2, arXiv:2405.04434.
- MoE (Mixture of Experts) ch. 08
- Replacing each dense MLP with many experts plus a router that activates a few per token, multiplying stored capacity while holding per-token compute nearly constant.
- outer product ch. 04
- kv⊤: a d×d rank-1 matrix that associates key direction k with value v. Linear attention adds these matrices to its state.
- prefill ch. 03
- Processing the whole prompt in parallel before generation starts. Compute-bound and quadratic in prompt length.
- PreNorm ch. 09
- Placing layer normalization before each sub-layer, leaving the residual stream itself un-normalized, the standard recipe, and the reason stream magnitude grows with depth.
- quantization-aware training (QAT) ch. 10
- Training (or fine-tuning) while simulating the low-precision number format used at serving time, so the model adapts to the rounding. K3: MXFP4 weights / MXFP8 activations from SFT onward.
- query / key / value ch. 02
- The three learned projections of attention: what am I looking for / what is stored here / the content itself. Color-coded this way throughout the explainer.
- residual stream ch. 01, 09
- The running vector for each token to which layers add their outputs. It carries accumulated information across depth; the term comes from Anthropic's interpretability work.
- RNN / recurrent state ch. 04
- A model that carries a fixed-size state forward step by step. Linear attention is attention re-derived as an RNN whose state is a matrix.
- router / load balancing / capacity factor ch. 08
- The small scoring network choosing top-k experts per token; the auxiliary pressure keeping traffic spread across experts; and the per-expert buffer size beyond which tokens overflow.
- SiTU ch. 08
- "Sigmoid Tanh Unit": the gated activation in K3's expert path (SiTU-GLU in the model card). Vendor-reported design detail; its lesson here is about kernel fusion, not the formula.
- softmax ch. 01, 02
- Exponentiate and normalize: turns a score vector into a probability distribution. Sharpens gaps between scores; see ch. 02's sharpness slider.
- state-space model / Mamba ch. 06
- The parallel lineage of fixed-state sequence models from which input-dependent decay gates entered the story. Gu & Dao, arXiv:2312.00752.
- token ch. 01
- A vocabulary chunk (word piece), the unit models read, predict, and measure context length in.