The Lab · Exercises and build-along videos
There are two tracks. The eight exercises recreate chapter demos in code, extend them, and ask you to produce a plot, table, or tool. The Karpathy track builds a GPT-style model from scratch and then reproduces GPT-2. You can follow either track or both.
Andrej Karpathy's Neural Networks: Zero to Hero starts with backpropagation on scalarscontext The first lecture builds micrograd, an automatic-differentiation engine in about 100 lines of Python. and ends with GPT-2. These two videos connect most directly to the explainer:
Watch "Let's build GPT" after chapter 02. "Let's reproduce GPT-2" fits after chapter 03, when you have seen batching, precision, and throughput.
Each exercise starts with a chapter's memory contract. Recreate its demo in code, change one part of the design, measure the result, and produce the requested plot, table, or tool. NumPy is enough until Lab 6. The source article also suggests a different study path.
Build: a ~40-line script that takes any architecture config (layers, KV heads, head dim, weight precision, KV-cache precision, context, batch) and prints its memory bill: weights, KV cache per token, cache at context, and, the interesting number, the crossover context where the cache becomes larger than the weights.
Extend: run it on GPT-2 small, Llama-2-7B, Llama-3-8B, and Kimi K3's published config (ch. 10 spec table; state any assumed KV-cache precision). Deliverable: one table, plus the crossover context for each model. Contract clause changed: none; the exercise makes the storage costs visible before you change a design.
Reproduce: write random unit key/value pairs into a d×d state additively; measure retrieval error as the count passes d:
import numpy as np
d, trials = 16, 6
errs = np.zeros(4 * d)
for _ in range(trials):
S, ks, vs = np.zeros((d, d)), [], []
for n in range(4 * d):
k = np.random.randn(d); k /= np.linalg.norm(k)
v = np.random.randn(d); v /= np.linalg.norm(v)
S += np.outer(k, v); ks.append(k); vs.append(v)
errs[n] += np.mean([np.linalg.norm(k_ @ S - v_)
for k_, v_ in zip(ks, vs)]) / trials
print(errs.round(2)) # near 0 while n << d, then grows without bound
You should see: error growing like √(n/d), roughly 0.8–1.0 by n=d, then past 1.0 and climbing without bound.
Extend, build a fullness gauge: the widget shows error after the fact;
build a meter that sees saturation coming. Track the state's effective rank
(from np.linalg.svd(S): singular-value energy, e.g. how many σ² it takes to reach 90%
of the total) as writes accumulate. Then repeat with orthogonal keys
(np.linalg.qr(np.random.randn(d, d))[0]) and watch the same gauge climb cleanly to d
and stop. Deliverable: one figure, two curves (retrieval error and effective rank
vs. writes) for random and orthogonal keys. Contract clause: the
"failure mode" row of chapter 04's contract becomes something you can measure
before retrieval error gets large.
Reproduce: store a value under a key, then store a different value under the same key; read it back both ways:
k = np.random.randn(16); k /= np.linalg.norm(k)
v1, v2 = np.random.randn(16), np.random.randn(16)
S_add = np.outer(k, v1) + np.outer(k, v2) # blind accumulation
S_del = np.outer(k, v1)
S_del += np.outer(k, v2 - S_del.T @ k) # delta: subtract what's there first
print(np.linalg.norm(k @ S_add - v2)) # large: v1+v2 mixed together
print(np.linalg.norm(k @ S_del - v2)) # ~0: cleanly replaced
Extend, a contradiction workload: generate a stream where 30% of writes update an existing key (a fact being revised) and 70% are fresh. Score both memories on "latest value wins": what fraction of reads return the current binding? Sweep β from 0.1 to 1.0 and find where partial trust beats full overwrite (hint: make some updates "noise" that should be ignored). Deliverable: accuracy vs. β curve. Contract clause: "eviction policy: none" becomes "targeted overwrite," with write strength as a tunable parameter.
Reproduce: ch. 06's topic shift: write topic A's associations, shift, write topic B's; compare stale-energy contamination under no decay, scalar decay at the boundary, and channel-wise decay.
Extend, a full policy grid: wrap one associative-memory class with pluggable
policies (add, delta, delta+scalar-gate,
delta+channel-gate) and run all four against three workloads: stable facts
(nothing should be forgotten), contradictions (Lab 2's stream), and topic shifts.
Score each cell. Deliverable: a 4×3 policy-by-workload table, your own miniature of
the ablation tables in the Gated DeltaNet paper. If one policy wins on every
workload, inspect the implementation and workload choices. Contract clause:
compare when keeping, correcting, and fading stored information each help.
Reproduce: implement a token-by-token recurrence (a Python loop of tiny matmuls) and a chunked version (reshape into chunks, big batched matmuls, carry state between chunks). Count operations; then time both on a GPU (or even NumPycontextNumPy shows the effect too, because the contrast (an interpreted Python loop versus vectorized C kernels) mirrors the GPU's serial-versus-parallel economics. The speedup depends on the array sizes and hardware; measure it rather than assuming a fixed range.). The chunked version can do more arithmetic and still finish faster.
Extend, find your hardware's chunk size: sweep C from 1 to T and plot both curves on one chart: total FLOPs (rises with C) and wall-clock (falls, bottoms out, then rises). The best chunk size can change between CPU and GPU. Deliverable: the two-curve plot with the fastest chunk size marked. Contract clause: none; compare operation count with measured run time.
Reproduce: a toy two-tier stack: recurrent (linear-attention) memory in most layers, one exact-attention layer periodically. Test on (a) verbatim copying of a random 20-token string seen long ago, and (b) associative recall (key→value pairs sprinkled through filler). Ablate the exact layer and watch copying collapse.
Extend, sweep the ration: vary the interleave, exact attention every k-th layer for k = 1, 2, 4, 8, ∞, and plot task accuracy against total memory cost. Somewhere on that curve is a knee; Kimi Linear shipped k=4 (3:1). Your toy won't reproduce their number, and that's the useful tradeoff: its location depends on the workload. Deliverable: the accuracy-versus-cost curve with the best tradeoff marked. Contract clause: choose how often the stack can afford exact attention.
Reproduce: stack L random residual blocks (h ← h + f(h)) with comparable increment norms; record hidden-state norm by depth and each increment's norm relative to the final state. See how that ratio changes as L grows.
Extend, instrument the real thing: load actual GPT-2 small
(transformers, ~500 MB) and register forward hooks on every block. For a real
sentence, log (a) residual-stream norm by depth and (b) each block's contribution norm relative to
the stream it joins. You are now measuring, in a real trained network, the dilution that ch. 09's
widget only sketched, including whatever surprises the trained weights have for the toy model's
assumptions. Deliverable: the norm-by-depth and contribution-share plots, plus one
sentence on where the real model differs from the equal-norm toy.
Contract clause: measure the depth-memory behavior in a trained model.
Take a frontier model card published after this explainer and fill out its six-memory contract: what plays the scratchpad, the ledger, the library, the depth stream; what precision; what interconnect assumptions. Note every cell the card leaves unstated, and every efficiency claim missing a scope (context length? batch? baseline? see ch. 10's checklist). Deliverable: one filled-in contract table and a list of the card's unanswered questions. The goal is to apply the same storage and retrieval questions to a model that the explainer has not covered.
These primary sources follow the chapter sequence:
Contents · Glossary · Videos by Andrej Karpathy, embedded with attribution; exercises adapted from the source article.