The Lab · Exercises and build-along videos

Build it yourself

There are two tracks. The eight exercises recreate chapter demos in code, extend them, and ask you to produce a plot, table, or tool. The Karpathy track builds a GPT-style model from scratch and then reproduces GPT-2. You can follow either track or both.

AThe Karpathy track

Andrej Karpathy's Neural Networks: Zero to Hero starts with backpropagation on scalarscontext The first lecture builds micrograd, an automatic-differentiation engine in about 100 lines of Python. and ends with GPT-2. These two videos connect most directly to the explainer:

Let's build GPT: from scratch, in code, spelled out. Video by Andrej Karpathy · builds chapter 01–02's machine, line by line · course repo · watch on YouTube ↗
Let's reproduce GPT-2 (124M) Video by Andrej Karpathy · trains the actual 124M model of chapter 01, with the systems details · build-nanogpt repo (step-by-step commits) · watch on YouTube ↗

Watch "Let's build GPT" after chapter 02. "Let's reproduce GPT-2" fits after chapter 03, when you have seen batching, precision, and throughput.

BThe exercises

Each exercise starts with a chapter's memory contract. Recreate its demo in code, change one part of the design, measure the result, and produce the requested plot, table, or tool. NumPy is enough until Lab 6. The source article also suggests a different study path.

Lab 0 · Calculate memory use (warm-up · ch. 03)

Build: a ~40-line script that takes any architecture config (layers, KV heads, head dim, weight precision, KV-cache precision, context, batch) and prints its memory bill: weights, KV cache per token, cache at context, and, the interesting number, the crossover context where the cache becomes larger than the weights.

Extend: run it on GPT-2 small, Llama-2-7B, Llama-3-8B, and Kimi K3's published config (ch. 10 spec table; state any assumed KV-cache precision). Deliverable: one table, plus the crossover context for each model. Contract clause changed: none; the exercise makes the storage costs visible before you change a design.

Lab 1 · Measure interference in fixed-size state (ch. 04's widget, then further)

Reproduce: write random unit key/value pairs into a d×d state additively; measure retrieval error as the count passes d:

import numpy as np
d, trials = 16, 6
errs = np.zeros(4 * d)
for _ in range(trials):
    S, ks, vs = np.zeros((d, d)), [], []
    for n in range(4 * d):
        k = np.random.randn(d); k /= np.linalg.norm(k)
        v = np.random.randn(d); v /= np.linalg.norm(v)
        S += np.outer(k, v); ks.append(k); vs.append(v)
        errs[n] += np.mean([np.linalg.norm(k_ @ S - v_)
                            for k_, v_ in zip(ks, vs)]) / trials
print(errs.round(2))   # near 0 while n << d, then grows without bound

You should see: error growing like √(n/d), roughly 0.8–1.0 by n=d, then past 1.0 and climbing without bound.

Extend, build a fullness gauge: the widget shows error after the fact; build a meter that sees saturation coming. Track the state's effective rank (from np.linalg.svd(S): singular-value energy, e.g. how many σ² it takes to reach 90% of the total) as writes accumulate. Then repeat with orthogonal keys (np.linalg.qr(np.random.randn(d, d))[0]) and watch the same gauge climb cleanly to d and stop. Deliverable: one figure, two curves (retrieval error and effective rank vs. writes) for random and orthogonal keys. Contract clause: the "failure mode" row of chapter 04's contract becomes something you can measure before retrieval error gets large.

Lab 2 · Compare additive and corrective updates (ch. 05's widget, then further)

Reproduce: store a value under a key, then store a different value under the same key; read it back both ways:

k = np.random.randn(16); k /= np.linalg.norm(k)
v1, v2 = np.random.randn(16), np.random.randn(16)

S_add = np.outer(k, v1) + np.outer(k, v2)          # blind accumulation
S_del = np.outer(k, v1)
S_del += np.outer(k, v2 - S_del.T @ k)             # delta: subtract what's there first

print(np.linalg.norm(k @ S_add - v2))  # large: v1+v2 mixed together
print(np.linalg.norm(k @ S_del - v2))  # ~0: cleanly replaced

Extend, a contradiction workload: generate a stream where 30% of writes update an existing key (a fact being revised) and 70% are fresh. Score both memories on "latest value wins": what fraction of reads return the current binding? Sweep β from 0.1 to 1.0 and find where partial trust beats full overwrite (hint: make some updates "noise" that should be ignored). Deliverable: accuracy vs. β curve. Contract clause: "eviction policy: none" becomes "targeted overwrite," with write strength as a tunable parameter.

Lab 3 · Compare ways to forget old information (ch. 06's widget, then further)

Reproduce: ch. 06's topic shift: write topic A's associations, shift, write topic B's; compare stale-energy contamination under no decay, scalar decay at the boundary, and channel-wise decay.

Extend, a full policy grid: wrap one associative-memory class with pluggable policies (add, delta, delta+scalar-gate, delta+channel-gate) and run all four against three workloads: stable facts (nothing should be forgotten), contradictions (Lab 2's stream), and topic shifts. Score each cell. Deliverable: a 4×3 policy-by-workload table, your own miniature of the ablation tables in the Gated DeltaNet paper. If one policy wins on every workload, inspect the implementation and workload choices. Contract clause: compare when keeping, correcting, and fading stored information each help.

Lab 4 · FLOPs are not seconds (the ch. 03 + 05 hardware lesson)

Reproduce: implement a token-by-token recurrence (a Python loop of tiny matmuls) and a chunked version (reshape into chunks, big batched matmuls, carry state between chunks). Count operations; then time both on a GPU (or even NumPycontextNumPy shows the effect too, because the contrast (an interpreted Python loop versus vectorized C kernels) mirrors the GPU's serial-versus-parallel economics. The speedup depends on the array sizes and hardware; measure it rather than assuming a fixed range.). The chunked version can do more arithmetic and still finish faster.

Extend, find your hardware's chunk size: sweep C from 1 to T and plot both curves on one chart: total FLOPs (rises with C) and wall-clock (falls, bottoms out, then rises). The best chunk size can change between CPU and GPU. Deliverable: the two-curve plot with the fastest chunk size marked. Contract clause: none; compare operation count with measured run time.

Lab 5 · Allocate a memory budget (ch. 07's argument, made quantitative)

Reproduce: a toy two-tier stack: recurrent (linear-attention) memory in most layers, one exact-attention layer periodically. Test on (a) verbatim copying of a random 20-token string seen long ago, and (b) associative recall (key→value pairs sprinkled through filler). Ablate the exact layer and watch copying collapse.

Extend, sweep the ration: vary the interleave, exact attention every k-th layer for k = 1, 2, 4, 8, ∞, and plot task accuracy against total memory cost. Somewhere on that curve is a knee; Kimi Linear shipped k=4 (3:1). Your toy won't reproduce their number, and that's the useful tradeoff: its location depends on the workload. Deliverable: the accuracy-versus-cost curve with the best tradeoff marked. Contract clause: choose how often the stack can afford exact attention.

Lab 6 · Depth telemetry on a real model (ch. 09, beyond the toy)

Reproduce: stack L random residual blocks (h ← h + f(h)) with comparable increment norms; record hidden-state norm by depth and each increment's norm relative to the final state. See how that ratio changes as L grows.

Extend, instrument the real thing: load actual GPT-2 small (transformers, ~500 MB) and register forward hooks on every block. For a real sentence, log (a) residual-stream norm by depth and (b) each block's contribution norm relative to the stream it joins. You are now measuring, in a real trained network, the dilution that ch. 09's widget only sketched, including whatever surprises the trained weights have for the toy model's assumptions. Deliverable: the norm-by-depth and contribution-share plots, plus one sentence on where the real model differs from the equal-norm toy. Contract clause: measure the depth-memory behavior in a trained model.

Lab 7 · Audit a model card (capstone · no code)

Take a frontier model card published after this explainer and fill out its six-memory contract: what plays the scratchpad, the ledger, the library, the depth stream; what precision; what interconnect assumptions. Note every cell the card leaves unstated, and every efficiency claim missing a scope (context length? batch? baseline? see ch. 10's checklist). Deliverable: one filled-in contract table and a list of the card's unanswered questions. The goal is to apply the same storage and retrieval questions to a model that the explainer has not covered.

CPaper reading order

These primary sources follow the chapter sequence:

  1. Transformers are RNNs (2020) · after ch. 04
  2. Linear Transformers Are Secretly Fast Weight Programmers (2021) · after ch. 05
  3. Parallelizing Linear Transformers with the Delta Rule (2024), alongside Yang's blog I–III · after ch. 05
  4. Gated Delta Networks (2024) · after ch. 06
  5. DeepSeek-V2 §2 (MLA) · after ch. 07
  6. Kimi Linear (2025) · after ch. 07
  7. Attention Residuals (2026) · after ch. 09
  8. Kimi K3 blog + model card · after ch. 10, with the checklist in hand

Contents · Glossary · Videos by Andrej Karpathy, embedded with attribution; exercises adapted from the source article.