A transformer language model, written from scratch in C,
generating in your browser.
~600 lines of hand-written C: RMSNorm, rotary embeddings, multi-head attention with a KV cache, a BPE tokenizer and a top-p sampler. Three self-trained models (an instruction-tuned one that follows a prompt, its TinyStories base and a byte-level Shakespeare model) are all int8-quantized and run by the same C engine compiled to WebAssembly. Pick one below. No libraries, no frameworks, no API calls.
loading 8 MB of int8 weights…The Instruct model wraps your text in a User: … Assistant: template and
writes a story to order. Try naming three words. Each passage ends itself: the model emits its
beginning-of-sequence token when the turn is done. Temperature 0 is greedy and deterministic; crank it
past 1.0 for chaos. Every token is a full forward pass through all six layers, in C, on your
CPU, and the models even swap tokenizers (byte-level vs 4096-vocab BPE) live in the engine.
How it works: the actual code
Every snippet below is really in the engine generating above. Follow one token through the machine: from a file of floats, through the forward pass, out the sampler. After that, see how the weights got trained in the first place.
The whole thing is one loop
A language model is a function from tokens so far to a probability for every possible next token. Generation just calls it repeatedly: predict, sample, append, repeat. While your prompt is still being fed in, the "next token" is forced to be your prompt; after that, the model is on its own.
while (pos < steps) {
float* logits = forward(t, token, pos); // the whole network, once
next = (pos < num_prompt_tokens - 1)
? prompt_tokens[pos + 1] // still feeding your prompt
: sample(s, logits); // now it's the model's turn
pos++;
if (next == 1) break; // it emitted BOS: it chose to stop
printf("%s", decode(tok, token, next)); // stream the piece out
token = next;
}A model is a file of floats
The entire model is one flat binary: a 28-byte header of seven ints, then every weight matrix as raw float32, in a fixed order. We mmap it and never copy: the "loaded model" is just pointers into the file.
typedef struct {
int dim; // width of the residual stream (192 here)
int hidden_dim; // feed-forward inner width (512)
int n_layers; // transformer blocks stacked (5)
int n_heads; // attention heads (6)
int n_kv_heads; // key/value heads (6, no grouping)
int vocab_size; // distinct tokens (259)
int seq_len; // max context length (256)
} Config;
fread(config, sizeof(Config), 1, file); // the 28-byte header...
*data = mmap(NULL, *file_size, PROT_READ, MAP_PRIVATE, *fd, 0); // ...the restLaying pointers over the blob
No deserialization. Each weight tensor is found by walking a pointer through the mapped file in the exact order the exporter wrote it. This ordering is the file format. One neat trick: if the header's vocab_size is positive, the final classifier is tied to the embedding table: the same matrix used twice.
w->token_embedding_table = ptr; ptr += vocab_size * dim;
w->rms_att_weight = ptr; ptr += n_layers * dim;
w->wq = ptr; ptr += n_layers * dim * dim;
w->wk = ptr; ptr += n_layers * dim * kv_dim;
// ...wv, wo, ffn norms, w1, w2, w3, final norm, in fixed order...
w->wcls = shared_weights ? w->token_embedding_table // tied: reuse it
: ptr; // untied: it followsOne vector, edited by every layer
The spine of a transformer is the residual stream: the token's embedding, which each sub-layer reads, computes a correction to and adds back. Information is never overwritten, only edited. That additive structure is why deep transformers train stably.
// x IS the model's running "thought"; start it as the token's embedding
memcpy(x, w->token_embedding_table + token * dim, dim * sizeof(float));
for (int l = 0; l < p->n_layers; l++) {
// attention sub-layer: compute a correction from context...
matmul(s->xb2, s->xb, w->wo + l*dim*dim, dim, dim);
for (int i = 0; i < dim; i++) x[i] += s->xb2[i]; // ...and ADD it
// feed-forward sub-layer: same pattern
for (int i = 0; i < dim; i++) x[i] += s->xb[i]; // never overwriteOne token in, one probability distribution out, looped.
Keeping the numbers sane
Before every sub-layer, the stream is rescaled to unit RMS and given a learned per-dimension gain. Five lines. Without it, activations drift in scale as they pass through dozens of matmuls and training falls apart.
void rmsnorm(float* o, float* x, float* weight, int size) {
float ss = 0.0f;
for (int j = 0; j < size; j++) ss += x[j] * x[j];
ss = 1.0f / sqrtf(ss / size + 1e-5f); // 1 / rms(x)
for (int j = 0; j < size; j++) o[j] = weight[j] * (ss * x[j]);
}Where all the time goes
Every projection in the network (queries, keys, values, outputs, both feed-forward wings, the classifier) is this one function. Each output element is a dot product of one weight row with the input. This loop is ~95% of the runtime; every LLM speed trick you've heard of (quantization, GPU kernels, batching) is fundamentally about making this faster.
// out = W @ x where W is (d, n) row-major, x is (n,)
void matmul(float* out, float* x, float* w, int n, int d) {
for (int i = 0; i < d; i++) {
float val = 0.0f;
float* wi = w + i * n; // row i of W
for (int j = 0; j < n; j++) val += wi[j] * x[j];
out[i] = val;
}
}Position as rotation
The network has no idea where a token sits; attention is order-blind. RoPE fixes this by rotating each adjacent pair of query/key dimensions by an angle proportional to the token's position (faster for low dims, slower for high: a spectrum of frequencies). The magic: after rotation, the dot product between a query and key depends only on their relative distance.
for (int i = 0; i < dim; i += 2) {
float freq = 1.0f / powf(10000.0f, (i % head_size) / (float)head_size);
float fcr = cosf(pos * freq), fci = sinf(pos * freq);
// rotate the pair (v0, v1) by angle pos*freq (complex multiplication)
float v0 = vec[i], v1 = vec[i + 1];
vec[i] = v0 * fcr - v1 * fci;
vec[i + 1] = v0 * fci + v1 * fcr;
}The only place tokens meet
Everything else operates on each position independently. Attention is where context flows. Each head scores the current token's query against every previous token's key, softmaxes the scores into weights, and takes that weighted blend of their values. Past keys/values sit in the KV cache.
And notice what's missing: there is no causal mask anywhere. A token physically cannot see the future, because the future isn't in the cache yet. The loop bound is the mask.
for (int t = 0; t <= pos; t++) { // every position so far
float* kt = key_cache + t*kv_dim + head_offset;
float score = 0.0f;
for (int i = 0; i < head_size; i++) score += q[i] * kt[i];
att[t] = score / sqrtf(head_size); // scaled dot product
}
softmax(att, pos + 1); // scores -> weights
for (int t = 0; t <= pos; t++) { // blend the values
float* vt = value_cache + t*kv_dim + head_offset;
for (int i = 0; i < head_size; i++) xb[i] += att[t] * vt[i];
}Where the knowledge lives
After attention gathers context, the feed-forward block does the per-token "thinking": expand to a wider hidden space through two parallel projections, gate one with the other using SiLU, contract back down. Most of the model's parameters (and most of what it knows) are in these three matrices per layer.
matmul(s->hb, s->xb, w->w1 + ..., dim, hidden_dim); // gate path
matmul(s->hb2, s->xb, w->w3 + ..., dim, hidden_dim); // up path
for (int i = 0; i < hidden_dim; i++) {
float val = s->hb[i];
val *= 1.0f / (1.0f + expf(-val)); // SiLU: x * sigmoid(x)
s->hb[i] = val * s->hb2[i]; // gated by the up path
}
matmul(s->xb, s->hb, w->w2 + ..., hidden_dim, dim); // back down259 opinions about what comes next
After the last layer: one final norm, then the residual stream is scored against every row of the (tied) embedding table. The result is a raw score, a logit, for each of the 259 tokens in the vocabulary. That vector is the model's entire output; everything after this is just deciding how to pick from it.
rmsnorm(x, x, w->rms_final_weight, dim);
matmul(s->logits, x, w->wcls, dim, p->vocab_size); // one score per token
return s->logits;One byte, one token
Real Llama uses 32,000 learned BPE merges. This model goes simpler: the vocabulary is just the 256 possible bytes (plus 3 specials). Nothing to train, and the C encode/decode machinery (including the full BPE merge loop) works unchanged; it simply finds nothing to merge. The price: one byte = one token, so 256 context tokens ≈ 256 characters. The TinyStories model (step 18) swaps in a real 4096-vocab BPE tokenizer: same C code, learned merges instead of none.
# id 0..2: <unk> <s> </s>, then one token per raw byte
def token_string(b):
if 32 <= b <= 126 or b in (9, 10, 13): # printable + whitespace
return bytes([b]) # the byte IS the token string
return f"<0x{b:02X}>".encode() # run.c already decodes these
tokens = [b"<unk>", b"<s>", b"</s>"] + [token_string(b) for b in range(256)]Rolling the dice, carefully
Temperature reshapes the distribution, then top-p cuts off the long tail of unlikely tokens before drawing. That's the whole difference between the temp-0 and temp-1.4 output in the demo above. Same model, same logits, different dice.
if (temperature == 0.0f) return sample_argmax(logits, n); // greedy
for (int q = 0; q < n; q++) logits[q] /= temperature; // reshape...
softmax(logits, n); // ...normalize...
float coin = random_f32(&rng_state); // ...and roll
return sample_topp(logits, n, topp, probindex, coin); // within the nucleusThe same model, written twice
The weights were trained by a PyTorch model that mirrors run.c exactly: same RMSNorm epsilon, same RoPE base and pairing convention, same tied classifier. Get any one of those wrong and the exported model produces fluent garbage: the file loads fine, the math runs fine and every number is subtly wrong. The two implementations agreeing is the correctness proof.
@dataclass
class ModelArgs:
dim: int = 192 # residual stream width
n_layers: int = 5
n_heads: int = 6
vocab_size: int = 259 # byte-level
hidden_dim: int = 512
max_seq_len: int = 256
norm_eps: float = 1e-5 # MUST match run.c's rmsnorm
# RoPE: base 10000, ADJACENT-pair rotation (i, i+1), run.c's
# convention, NOT HuggingFace's rotate-half. One wrong knob = garbage.Watching it learn, then memorize
2.26M parameters, ~1.1M tokens of Shakespeare, AdamW, six minutes on a laptop GPU. Loss starts at 5.56 (exactly ln(259), the loss of pure guessing) and the curves below are the real run. Around iteration 750 the validation loss bottoms out and turns back up while training loss keeps falling: the model has started memorizing its corpus instead of learning English. The shipped checkpoint is the one from the bottom of the valley.
The real loss curves from this model's training run. Divergence = overfitting, live.
The handshake between two worlds
Export is just serialization discipline: write the 7-int header, then dump every tensor in exactly the order step 03's pointer walk reads them. PyTorch stores linear weights as (out, in) row-major, the very layout run.c's matmul indexes, so there's not even a transpose. Train in Python, run anywhere C runs. Which, via WebAssembly, includes the page you're reading.
f.write(struct.pack("<7i", dim, hidden_dim, n_layers,
n_heads, n_kv_heads, vocab_size, seq_len))
serialize_fp32(f, model.tok_embeddings.weight)
for layer in model.layers: serialize_fp32(f, layer.attention_norm.weight)
for layer in model.layers: serialize_fp32(f, layer.attention.wq.weight)
# ...wk, wv, wo, ffn norms, w1, w2, w3, final norm:
# the exact order memory_map_weights() walks. This order IS the format.Four bytes into one
Step 06 said matmul is where all the time goes, but it's not the multiplies, it's streaming the weight bytes from memory. So: store each weight as an int8 plus one fp32 scale per group of 64, quantize the activations on the fly and run the inner loop in pure integer arithmetic. 4× less memory traffic, 4× smaller file. What stays fp32: the norms, the KV cache, everything between ops: precision where it's cheap, bytes saved where it's expensive.
This isn't hypothetical: the engine generating above is runq.c, and the 2.4 MB you downloaded (instead of 9.1) is the quantized checkpoint.
// activations -> int8, one scale per group, right before each matmul
float scale = wmax / 127.0f;
qx->q[i] = (int8_t)roundf(x[i] / scale);
// the payoff: the hot loop is integer multiplies with an int32 accumulator
for (int j = 0; j <= n - GS; j += GS) {
int32_t ival = 0;
for (int k = 0; k < GS; k++)
ival += ((int32_t)x->q[j + k]) * ((int32_t)w->q[in + j + k]);
val += ((float)ival) * w->s[(in + j) / GS] * x->s[j / GS]; // scales, once
}The bug that spoke in tongues
First int8 run of a dim-288 model: fluent English with random exotic tokens sprinkled in: "there was a little girl namedoptsa". Cause: 288 isn't divisible by the group size 64, so quantization groups straddled matrix rows and every odd row read its neighbour's scales. Half the vocabulary's logits corrupted, half perfect: a model that's simultaneously right and wrong. Our Shakespeare model (dims 192/512) masked the bug completely.
The lesson quantized models teach: they fail eloquently. Nothing crashes; the text just gets strange. The fix is one invariant (groups must align with rows), enforced twice:
# export.py: shrink the group size until it divides every row length
while cfg["dim"] % gs != 0 or cfg["hidden_dim"] % gs != 0:
gs //= 2
# runq.c: and refuse loudly rather than generate beautifully wrong text
# if (config->dim % GS != 0 || config->hidden_dim % GS != 0) exit(1);From letters to words
Try the 📖 TinyStories model at the top. It reads far better than Shakespeare, for two
reasons. First, a real byte-pair-encoding tokenizer, trained from scratch: start from the
256 bytes and repeatedly fuse the most common adjacent pair into a new token, to a vocab of 4096.
The first merges it learns on children's stories are he, t,
a, the: the statistical skeleton of English. Now "Once upon a
time" is ~4 tokens, not 16, so the same 256-token window reaches four times deeper into a story.
And, the recurring theme of this whole project, run.c needed zero changes. Its encode()
was already a greedy "merge the highest-scored pair" loop, which is exactly BPE. We just write the
learned merges into the same tokenizer file with score = -rank, and the C engine
reproduces our tokenization for free.
for i in range(n_merges):
pairs = Counter() # count every adjacent pair
for word, freq in corpus.items():
for a, b in zip(word, word[1:]):
pairs[(a, b)] += freq
(a, b), _ = pairs.most_common(1)[0] # the most frequent pair...
merges.append((a, b)) # ...becomes one new token
corpus = {merge(w, a, b, BASE + i): f for w, f in corpus.items()}Seven million parameters, a real corpus
Second reason: scale. The Shakespeare model saw ~1 MB of text; this one trained on
~300 MB of TinyStories: 79 million tokens, tokenized once into a memory-mapped
uint16 array so the GPU never waits on Python. The model itself grew to
7.2M parameters (dim 288, 6 layers), the shape of Karpathy's stories15M but with our
leaner 4096 vocab. One epoch is ~4,800 steps, so a few thousand steps of training never even
finishes a full pass, the opposite of the Shakespeare model, which memorized its tiny corpus.
The result is what you're reading: real words, whole sentences, little arcs with a beginning and
a "happily ever after".
# pre-tokenized corpus, memory-mapped; sample random 256-token windows
data = np.memmap("tinystories_train.bin", dtype=np.uint16, mode="r")
ix = np.random.randint(0, len(data) - seq_len - 1, size=batch)
x = np.stack([data[i : i+seq_len ] for i in ix])
y = np.stack([data[i + 1 : i+seq_len + 1] for i in ix]) # targets = inputs shifted by oneTeaching it to follow, not just continue
The base model continues text; the 💬 Instruct model at the top follows an instruction. Same 7M weights, further trained ("supervised fine-tuning") on a chat template built from the stories themselves: pull content words out of a story, and the instruction becomes "write a story using these words," with the story as the target answer:
User: Write a story using the words: dragon, cake, brave. Assistant: <the story></s>
The one genuinely new idea is loss masking. We do not want the model learning to
generate the instruction, only the response. So every target position inside the prompt is set to
-100, which PyTorch's cross-entropy silently ignores. The gradient only ever flows
from the answer. And, the theme of this whole project once more, run.c is unchanged: the
template is plain text, and BOS is still the stop token, so the model just learns to end its turn.
full = prompt_ids + response_ids + [BOS] # BOS = "end of turn"
x, y = full[:-1], full[1:]
# only positions predicting a RESPONSE token count toward the loss:
for i in range(len(prompt_ids) - 1):
y[i] = -100 # cross_entropy ignores these
# ...later, unchanged: loss = F.cross_entropy(logits, y) # ignore_index=-100What 7M parameters can and can't do
Set expectations: this is not an assistant. It's a demonstration that the mechanism works. Ask for a story and it obeys the format and theme every time: "about a robot" reliably gets you a robot. Injecting specific requested words is harder: common ones (dog, park, friend) usually land, rare ones often don't. About 24% of random content words make it in (measured on held-out prompts). It has no facts, no reasoning, no memory across turns; it knows only the small world of TinyStories.
Two things it does not use, that real instruct models do: LoRA/PEFT (we full-fine-tune all 7M weights because at this size it's free; large models train a small adapter instead), and RLHF/DPO (a second preference-alignment stage after SFT). SFT is step one, where instruction-following comes from. That 24% is the number the next phase attacks.
Teaching taste without a teacher
SFT taught the model to answer. Alignment teaches it which answer is better. Real RLHF collects pairs where a human picked the better of two responses; we can't hand-label thousands, so we use a programmatic judge, which makes this RLAIF, but the optimizer downstream is identical.
The judge is Phase 6's weakness itself: did the story use the requested words? For each prompt we sample several completions from the SFT model, then pair the one using the most requested words (chosen) against the one using the fewest (rejected). The preferences are the model's own outputs, sorted by a rule it will now be trained to satisfy.
completions = sample_k(model, prompt_ids, k=8) # on-policy: the SFT model's own
scored = [(sum(w in text.lower() for w in words), text) # reward = # requested words present
for text in completions]
scored.sort()
rejected, chosen = scored[0][1], scored[-1][1] # worst vs best: a preference pairTwo models, one nudge
Classic RLHF trains a reward model then optimizes it with RL (PPO). DPO skips both: it optimizes the preference pairs directly. Keep two copies of the SFT model (a trainable policy and a frozen reference) and push the policy to prefer chosen over rejected, but only beyond what the reference already prefers. One line of loss, no reward model, no RL loop.
The result, both numbers measured identically on the same held-out prompts: word-inclusion went from 22% (SFT) to 45% (DPO), roughly doubled; the model learned to actually use the words it's given. The cost, visible in its output, is the alignment tax: reward-hacking that trades fluency for score. Switch between 💬 Instruct and 🎯 Aligned above and give both the same three words. The difference is DPO.
pol = logp(policy, chosen) - logp(policy, rejected) # does the policy prefer chosen?
ref = logp(ref, chosen) - logp(ref, rejected) # did the frozen SFT model already?
loss = -logsigmoid(beta * (pol - ref)).mean() # reward only the NEW preferenceThe machinery DPO threw away
DPO was invented to replace this. Before it, alignment meant PPO: reinforcement learning, the InstructGPT recipe. Same reward, same frozen reference, but instead of one offline loss over preference pairs, PPO runs a full RL loop: the policy generates responses on-policy, a value head (critic) scores each state, GAE turns rewards into advantages and a clipped surrogate objective nudges the policy, clipping the probability ratio so no update moves too far (the "proximal" in PPO), with a per-token KL penalty back to the reference.
Five moving parts (rollouts, reward, critic, GAE, clipped+KL loss) regenerated every iteration, versus DPO's one loss over a fixed dataset. That extra machinery is the comparison.
ratio = exp(new_logp - old_logp) # how far the policy moved, per token
clipped = clamp(ratio, 1-eps, 1+eps) * adv # don't trust a big move too far
loss = -min(ratio*adv, clipped).mean() \ # clipped policy gradient
+ vf_coef*value_loss - ent_coef*entropy # + train the critic + reward explorationTwo roads, same destination
Both start from the SFT model and chase the same reward. On the held-out word-inclusion test: SFT 22% → DPO 45% → PPO ~22%. That's the honest result: our PPO didn't beat the baseline. Not because PPO can't (it's the algorithm behind ChatGPT) but because it's finicky. Stable settings (higher KL penalty) barely moved; aggressive ones (higher LR, lower KL) destabilized: entropy blew up and reward fell. Across three configs, none cleanly beat DPO's robust first-try doubling.
That contrast is the lesson: PPO's high-variance policy gradient vs DPO's low-variance contrastive loss is exactly why DPO became the default for straightforward preference alignment. And the cost gap seals it. PPO needs a critic, GAE, reward shaping, KL/clip tuning and fresh rollouts every iteration (minutes of generation), for a result DPO reached with one loss over a static file. Switch between 🎯 Aligned (DPO) and 🕹️ PPO on the same words. DPO reliably works them in; our PPO, like SFT, often doesn't.
The piece both toys skipped
Why did PPO stall? Two missing pieces, and this phase adds them. First, a real reward model, the middle stage of RLHF that DPO folds away and our PPO faked with the raw rule. We train it on the same preference pairs DPO used (Phase 7's file), with the Bradley-Terry loss: score the chosen response above the rejected one. The result is a dense, smooth reward for any response, a far better gradient signal than the rule's coarse 0 / ⅓ / ⅔ / 1.
Honest caveat, visible at this scale: our reward model only reaches ~61% pairwise accuracy; a 7M linear-head probe on ~1k pairs is a weak judge. RL is only ever as good as its reward model, and ours is the weak link.
r_chosen = reward_model(prompt + chosen) # scalar score, mean-pooled hidden
r_rejected = reward_model(prompt + rejected)
loss = -logsigmoid(r_chosen - r_rejected).mean() # chosen should score higherRL without the machinery
Second missing piece: a low-variance estimator. RLOO (REINFORCE Leave-One-Out) drops PPO's critic, GAE and clipping entirely. For each prompt it draws k samples and scores them with the reward model; the baseline for each sample is simply the mean of the others. That leave-one-out baseline is the whole variance-reduction trick: no value network, no GAE, no clip. One REINFORCE update per rollout, KL-tethered to the reference.
First attempt, against the ~61% reward model above. RLOO's machinery worked perfectly: the reward-model score tripled (−0.6 → +2.4) and KL to the reference exploded 9× as the policy chased it. But word-inclusion barely budged (22% → 24%). The policy hacked the weak reward model, exploiting its errors instead of using the words. That's reward-model overoptimization (Goodhart's law: "when a measure becomes a target, it ceases to be a good measure"), the central pitfall of RLHF. The reward model, it turned out, didn't even score a 2-word story above a 1-word one, so "maximize it" didn't mean "use the words."
reward = reward_model(samples) - kl_coef * kl_to_ref # k samples for this prompt
baseline = (reward.sum() - reward) / (k - 1) # mean of the OTHER k-1
advantage = reward - baseline # unbiased, low-variance
loss = -(advantage.detach() * logp_policy).mean() # plain REINFORCE, no clipFix the reward, and it climbs
RL is only as good as its reward model, so we built a better one. The weak RM never learned to grade because its pairs were all-or-none (chosen had every word, rejected had none); it never saw "2 words vs 1." The fix is graded pairs, made for free from the corpus: hold the story fixed and vary the request: ask for words the story happens to contain (3/3 satisfied) vs words it doesn't (1/3). Now the reward model learns to rank by how many requested words appear, and on real generations it does: 2-word stories now score above 1-word above 0.
Same RLOO, same everything; only the reward model changed. And the identical algorithm that hacked now climbs: word-inclusion 22% → 32%, with KL bounded (~4, not 9). It climbed the reward in-distribution, genuinely using more words. The full arc: SFT 22% → PPO 22% → RLOO/weak-RM 24% (hacked) → RLOO/strong-RM 32% (climbed) → DPO 45%. RL from a reward model works, when the reward model is good. And DPO's quiet advantage endures: no reward model to get right in the first place. 🎁 RLOO above is this model. Switch it against 🕹️ PPO on the same words.
# same story A; make c of the 3 requested words ones that A actually contains
request(c) = (c words drawn from A) + (3 - c words NOT in A)
chosen = request(c_hi) # more requested words present
rejected = request(c_lo) # fewer present -> the RM learns to grade by countEvery benchmark has a floor
"45%" invites an obvious complaint: fewer than half the requested words show up, and that's not a good assistant. Fair. But before you accept or dismiss a number, ask what it scores when the model ignores the instruction entirely. TinyStories has a small, repetitive vocabulary, so a story that never read the request still contains ~11% of the requested words by pure coincidence. That's the floor, and it takes one line to measure: score each story against a different prompt's words.
Subtract it and the picture inverts. SFT scores 19.6% against a 10.4% floor: only +9 points of real instruction-following; a third of its headline was luck. DPO scores 43.3% against 11.7%, +32 points. So the honest comparison was never 22% → 45% (a 2× bump); it's 9 → 32, about 3.4×. The raw numbers were understating alignment.
The metric is also harsher than it sounds. The requested words are picked to be distinctive, which means 78% of them are rare. Split it out: on common words DPO hits 60%; on rare words 39%, up from SFT's 13%, nearly 3×. The headline is dominated by the hardest version of the task, and a 7M model with a 4096-token vocabulary may not reliably spell a rare word at all. Some of that miss rate is a capacity ceiling, not an alignment failure.
And one caveat aimed at ourselves: at n=240 the standard error is ±3 points. DPO's 45% vs RLOO's 32% is a real gap (~4 s.e.), but "SFT 22% vs PPO 22%" is just noise agreeing with noise. Read no precision into it. The lesson outlives the number: know your floor, report the lift, and know which part of the task your metric is actually measuring.
# real: does story i contain the words asked of story i?
hit = sum(w in gens[i] for w in reqs[i])
# floor: does story i contain the words asked of some OTHER story?
chance = sum(w in gens[i] for w in reqs[(i + 7) % n]) # ~11%, pure coincidence
lift = hit/tot - chance/tot # <- the only number that means anything