Chapter 08 · Activating only some parameters
The previous chapters focused on memory for the token sequence. Mixture of Experts (MoE) changes how the model uses its learned parameters. Kimi K3 stores 2.8 trillion parameters but activates only some of them for each token. That distinction matters when comparing its size or cost with a dense model.
A Transformer block has attention followed by an MLP. In large models, the MLPs holdcontext In GPT-2 small, each block's MLP has about 4.7M parameters (768→3072→768), compared with about 2.4M in its attention. That is roughly two-thirds of the block's parameters. most of the parameters. A dense model uses the full MLP for every token. MoE replaces it with many MLPs ("experts") and a learned router that selects a few for each token. The total parameter count can then grow faster than per-token computation. The lineage runs from Shazeer et al.'s sparsely-gated MoE (2017) through Switch Transformers (2021) to essentially every 2025-26 frontier model, the trend Raschka's architecture survey documents model by model.
The router is a small scoring function: for each token it ranks all experts and keeps the top-kcontext Selecting 16 of 896 small experts gives the router many possible combinations: about 7×10³³. More, smaller experts allow finer selection but also create more routing decisions and scattered memory traffic.. The selected outputs are combined as follows:
The balancing loss below exists only at training time, to shape what the router learns; at inference the router simply scores and routes.
Routing creates two recurring problems:
Shared experts, the always-on lane in Plate 8·A, hold common knowledge every token needs, so the routed experts can focus on less common patternscontextThe shared/routed split was popularized by DeepSeekMoE (2024): isolate the common stuff (function words, basic syntax) in always-on experts so the routed ones aren't each forced to relearn it. The same K3 model card lists two shared experts.. The demo shows how uneven routing affects capacity:
No one assigns an expert a subject such as medicine, French, or Python. Training shapes the experts and the router together, sorting tokens according to patterns that help the model predict what comes next.
Mixtral's eight experts offer a useful example. Its researchers found no clear division by topic. Instead, some choices tracked syntax and surface form: indentation tokens in code tended to go to the same experts, while Python's "self" and English "Question" could go to the same one (Jiang et al., 2024). That is what Mixtral's router learned; it does not tell us how K3's 896 experts divide their work.
In a dense model, every token passes through the entire feed-forward map. MoE can hold more feed-forward parameters while running only a small selection of them for each token. The router makes that selection, and the same experts can serve many tokens. How well the model predicts tokens depends on whether training finds a useful way to share that work.
The GPT-2 paper and K3 model card support two different parameter ratios:
A dense GPT-2 uses essentially all its parameters on every token; K3 consults about 3.7% of its parameterscontextThat fraction is a design dial, and it has been falling: Mixtral (2023) activated roughly a quarter of its weights per token (2 of 8 experts); DeepSeek-V3 (2024) about 5.5% (37B of 671B); K3 is at 3.7%. Total capacity is growing faster than per-token compute in these examples. per token. The activated-parameter ratio is about 839×, though it still does not directly measure compute or speed: attention structure, precision, and hardware utilization also matter. The source article makes the distinction this way:
Adapted from the source article · @waterloo_intern
Never compare MoE total parameters with dense-model parameters and call the result "per-token compute."
"22580: From GPT2 to Kimi3, Explained" (lightly reworded; the italicized rule is the article's).
Sparse activation saves arithmetic but creates routing and placement costs:
| Problem | What it is |
|---|---|
| Router balance | Popular experts overload the devices hosting them while others idle |
| All-to-all communication | Tokens must physically travel to the accelerators where their experts live, every MoE layer, both directions |
| Memory footprint | All 2.8T weights must be resident somewhere: sparsity saves compute, not storage |
| Batching | Different tokens choose different experts, fragmenting nice dense batches |
| Numerical stability | Very sparse routing changes optimization behavior at scale |
List adapted from the source article.
K3's answers, per Moonshot's technical blog (vendor-reported; independent reproduction at this scale doesn't exist): a Stable LatentMoE framework, in which token representations are compressed into a latent space before expert computation and projected back aftercontext Chapter 07 used learned compression for cached K/V data. Here the reported compression acts on token representations around expert computation. (cutting per-expert FLOPs and communication volume), and a SiTU activation ("Sigmoid Tanh Unit", appearing as SiTU-GLU in the model card) inside the expert path.
An activation function with a low operation count can still run slowly if it cannot be fused with neighboring GPU operations. Extra launches, memory transfers, and synchronization can outweigh the arithmetic savings. FLOP counts alone do not include those costs.
What this chapter established
Go deeper
Contents · Glossary · K3 figures from the official model card; blog-only claims labeled vendor-reported.