Chapter 08 · Activating only some parameters

Mixture of Experts

The previous chapters focused on memory for the token sequence. Mixture of Experts (MoE) changes how the model uses its learned parameters. Kimi K3 stores 2.8 trillion parameters but activates only some of them for each token. That distinction matters when comparing its size or cost with a dense model.

Memory contract · sparse Mixture of Experts (parameter memory)
Stored
All 2.8T weights, resident across many accelerators; capacity is a library
Consulted / token
16 of 896 routed experts + 2 shared experts; activated total per the model card: ~104B (a figure that includes attention and dense parts too, not experts alone)
Eviction policy
None needed; the analog is routing: choosing which capacity to consult, per token, per layer
Failure mode
Systems complexity: load imbalance, all-to-all communication, and a memory footprint that must live somewhere

8.1Where the parameters actually are

A Transformer block has attention followed by an MLP. In large models, the MLPs holdcontext In GPT-2 small, each block's MLP has about 4.7M parameters (768→3072→768), compared with about 2.4M in its attention. That is roughly two-thirds of the block's parameters. most of the parameters. A dense model uses the full MLP for every token. MoE replaces it with many MLPs ("experts") and a learned router that selects a few for each token. The total parameter count can then grow faster than per-token computation. The lineage runs from Shazeer et al.'s sparsely-gated MoE (2017) through Switch Transformers (2021) to essentially every 2025-26 frontier model, the trend Raschka's architecture survey documents model by model.

Plate 8·AOne MoE layer in Kimi K3's configuration
token routerscores all 896 expert #17 expert #841 … 16 selected of 896 … 880 experts idle 2 shared expertsevery token, always Σ out
Expert counts from the Kimi K3 model card (896 routed, 16 selected, 2 shared), verified 2026-08-15. Weighted combination of expert outputs follows the standard MoE formulation.

8.2How routing works

The router is a small scoring function: for each token it ranks all experts and keeps the top-kcontext Selecting 16 of 896 small experts gives the router many possible combinations: about 7×10³³. More, smaller experts allow finer selection but also create more routing decisions and scattered memory traffic.. The selected outputs are combined as follows:

y = Σe ∈ top-k we · Experte(x)    w = softmax of the k winning router scores

The balancing loss below exists only at training time, to shape what the router learns; at inference the router simply scores and routes.

Routing creates two recurring problems:

Shared experts, the always-on lane in Plate 8·A, hold common knowledge every token needs, so the routed experts can focus on less common patternscontextThe shared/routed split was popularized by DeepSeekMoE (2024): isolate the common stuff (function words, basic syntax) in always-on experts so the routed ones aren't each forced to relearn it. The same K3 model card lists two shared experts.. The demo shows how uneven routing affects capacity:

Router playgroundInteractive

200 tokens stream into an 8-expert layer (top-1 routing for clarity). "Natural" token preferences are skewed: some experts are just more useful. Raise the balancing pressure and watch utilization even out; the dashed line is each expert's capacity.

Toy model: preference skew is a Zipf-like distribution, balancing pressure blends it toward uniform, a stand-in for the auxiliary-loss effect, not a simulation of any real router. Real systems route top-k (K3: top-16 of 896) and handle overflow in hardware-specific ways.

What an "expert" represents

No one assigns an expert a subject such as medicine, French, or Python. Training shapes the experts and the router together, sorting tokens according to patterns that help the model predict what comes next.

Mixtral's eight experts offer a useful example. Its researchers found no clear division by topic. Instead, some choices tracked syntax and surface form: indentation tokens in code tended to go to the same experts, while Python's "self" and English "Question" could go to the same one (Jiang et al., 2024). That is what Mixtral's router learned; it does not tell us how K3's 896 experts divide their work.

In a dense model, every token passes through the entire feed-forward map. MoE can hold more feed-forward parameters while running only a small selection of them for each token. The router makes that selection, and the same experts can serve many tokens. How well the model predicts tokens depends on whether training finds a useful way to share that work.

8.3What the 22,580× comparison measures

The GPT-2 paper and K3 model card support two different parameter ratios:

total: 2.8 × 10¹²124 × 10⁶ ≈ 22,580×     but     activated: 104 × 10⁹124 × 10⁶ ≈ 839×

A dense GPT-2 uses essentially all its parameters on every token; K3 consults about 3.7% of its parameterscontextThat fraction is a design dial, and it has been falling: Mixtral (2023) activated roughly a quarter of its weights per token (2 of 8 experts); DeepSeek-V3 (2024) about 5.5% (37B of 671B); K3 is at 3.7%. Total capacity is growing faster than per-token compute in these examples. per token. The activated-parameter ratio is about 839×, though it still does not directly measure compute or speed: attention structure, precision, and hardware utilization also matter. The source article makes the distinction this way:

Adapted from the source article · @waterloo_intern

Never compare MoE total parameters with dense-model parameters and call the result "per-token compute."

"22580: From GPT2 to Kimi3, Explained" (lightly reworded; the italicized rule is the article's).

8.4The systems cost of sparse activation

Sparse activation saves arithmetic but creates routing and placement costs:

ProblemWhat it is
Router balancePopular experts overload the devices hosting them while others idle
All-to-all communicationTokens must physically travel to the accelerators where their experts live, every MoE layer, both directions
Memory footprintAll 2.8T weights must be resident somewhere: sparsity saves compute, not storage
BatchingDifferent tokens choose different experts, fragmenting nice dense batches
Numerical stabilityVery sparse routing changes optimization behavior at scale

List adapted from the source article.

K3's answers, per Moonshot's technical blog (vendor-reported; independent reproduction at this scale doesn't exist): a Stable LatentMoE framework, in which token representations are compressed into a latent space before expert computation and projected back aftercontext Chapter 07 used learned compression for cached K/V data. Here the reported compression acts on token representations around expert computation. (cutting per-expert FLOPs and communication volume), and a SiTU activation ("Sigmoid Tanh Unit", appearing as SiTU-GLU in the model card) inside the expert path.

Why kernel implementation matters

An activation function with a low operation count can still run slowly if it cannot be fused with neighboring GPU operations. Extra launches, memory transfers, and synchronization can outweigh the arithmetic savings. FLOP counts alone do not include those costs.

Source spotlight: the MoE explainers this chapter compresses
Grootendorst's Visual Guide & Hugging Face's MoE Explained
Maarten Grootendorst's "A Visual Guide to Mixture of Experts" covers routing, load balancing, and expert capacity in more than 50 illustrations. Hugging Face's "Mixture of Experts Explained" adds the engineering history (Switch, expert parallelism) and training practicalities.

What this chapter established

Go deeper

Contents · Glossary · K3 figures from the official model card; blog-only claims labeled vendor-reported.