Chris Hay takes a transformer apart by hand on a locally-running Gemma 3 4B (no GPU, no cloud) to answer one question: how can a pile of frozen weights be a queryable database? He builds a packed knowledge store and a hand decoder from scratch — then finds the surprise. The model packs facts on top of each other but never unpacks them. It has no decoder and doesn't need one: it reads each fact by address, and the address is relation + entity, assembled progressively across layers ~23–27. He closes by hand-building an FFN as a literal key–value store and writing a brand-new fact into Gemma's own memory — no training — which the model then reads natively on its own forward pass.
OSLO.Hay opens on the payoff demo. Gemma 3 4B is loaded and verified clean of three invented facts — capital of Zealandia → Oslo, currency of Qataria → Yen, language of Vornhol → Welsh — and those values are hand-injected into slots 102039, 102038, 102037. No training. When asked, the model looks up those slots itself and answers correctly on its own forward pass.
To understand why that works, he backs up to representation. Nothing is stored as text. Slot 102039 holds six numbers — coordinates on a map — and “direction” just means the line from the origin to that point.
One fact per slot would exhaust the model instantly; there are millions of facts and finite slots. So the model packs: take the capital, currency, and language directions and add them up. Three facts, one vector, landing somewhere new in the space.
The code is two functions. pack is the addition. decode is where the work is — a match-and-peel algorithm: loop the known directions, find the loudest, subtract it out, repeat until only noise is left. Run it and the packed vector yields currency first (loudest), then capital, then language, then noise. It scales to real-world atlases packed into a single slot. This is superposition; the only difference from the model is that Hay hand-built both the pack and the decode.
Does the model carry a reader that peels facts apart? He tests it with three conditions on a made-up country packed with capital → Cairo, currency → Rand, language → Tamil, injected at layer 20 — early, giving the model every chance:
The conclusion is blunt: the model has no decoder. Which raises the real question — if its own weights are full of packed relations and it can't unpack them, how is it reading them at all?
Before answering, a definitional aside worth pinning down. Monosemantic = one fact per slot, nothing shared, clean. Polysemantic = many facts sharing directions, stacked on top of each other — and that's the one needing a decoder.
To check whether the model packs, he runs a simple linear reader inside the network. Layer 26 is chosen deliberately: tracing “the capital of France is,” Paris hasn't emerged at layer 20 or 22, starts emerging around 24, is locked in by 25, and is clean by 26. At L26 he probes three things off one residual vector — the value (the model's own unembed readout), the relation, and the entity (which of 65 places). Result: value ~0.75, relation picked up straight away, and the entity identified correctly — Hay reports it as a top-5 hit out of 65 and “spot on,” which against chance is minuscule.
So three different facts sit in one vector, all linearly readable, all polysemantically packed. Packing isn't hiding anything. What the model lacks isn't storage — storage is fine — it's demixing. And it doesn't need demixing, because it isn't reading the packed channel by peeling. It's reading each fact by an address. The three probes were linear: straight-line lookups. Reading is a lookup, and it's free.
Reading isn't all you want from a store of facts — sometimes you want to compute over them. He packs eight facts into shared slots on a toy model, hands them to one linear reader, and asks four jobs of it:
| Job | One linear reader (lookup) |
|---|---|
| Read the fact back out | ✅ perfect |
| Count | ✅ |
| Majority | ✅ |
| Parity / XOR | ❌ |
The parity failure is not a new finding — it's Minsky & Papert, 1969, one of the oldest limitations in the book, and Hay says so explicitly. What matters is the inference drawn from it: on a packed store of facts, the wall is not packed-vs-unpacked, it's lookup-vs-computation. The bits were always there; a lookup simply can't do XOR. Add one layer of real computation back in and parity returns.
"The model is packing to store, but it's addressing to read."
To make that concrete he queries “the capital of Australia” — famous-but-wrong-answer bait. At layer 20, nothing: neither Canberra nor Sydney. At layer 24 a fuzzy resolution appears — Sydney, the larger and more famous city, leads while Canberra sits low. By layer 26 Canberra is ahead, and by 28 it's locked. USA/Washington, Canada/Ottawa, Turkey/Ankara and Brazil/Brasília all behave the same way.
The model isn't looking in one place. It resolves layer by layer: a fuzzy search first, an exact search on the relation, and then it snaps in around L25–L27 — the fact band. This is why a fact can't be injected just anywhere. Put it in late and the model won't buy it; it has to be built at the right point, in the right directions, so it answers across the query variations.
The confirming experiment: for “the capital of France is Paris,” the FFN writes in the Paris direction happen around L23, L25 and L27. Zero out the last-position FFN writes between layers 23 and 28 and the answer collapses entirely — because the lookup is what writes the value into the residual stream.
Once the address is assembled — relation + entity, e.g. capital + France — the lookup fires and writes the value into the residual stream. From there the later layers (L28 onward, the output layers) simply read it as it rides along. The model does the pack on the write side, but on the read side it never decodes. That's the mechanism that lets everything be packed into one vector while the model stays a pure lookup with no demixer.
He trains a tiny probe at ~layer 10 to read one thing off the middle of the model: what relation is this? Trained on three words only — capital, currency, language — it nails them. Then he tests it on words it has never seen: seat, money, tongue. It generalizes. It knows seat means capital in the context of capitals.
So the address is semantic, not lexical. This is consistent with his earlier LARQL work: synonyms and antonyms are themselves connected by relationships at the earlier word layers — these relations don't only live in the fact layers.
The closing question from the top: could an FFN be built by hand? Yes.
There were two ways to pack. Superposition — everything loaded into a single channel, requiring a demixer to read, which the model can't do because it only has reads and lookups over the residual stream. And the model's way — a key–value store, which is all an FFN is: one row of the input matrix is an address (the spot that matches the slot); the matching row of the output matrix is the answer. Feed an address, the neuron fires, it writes its value. One matrix multiply. No demixing.
He runs exactly that: a hand-built FFN, W_in six rows × 24 dimensions. Each input row is the address a neuron detects; each output row is the value it writes. Six facts planted across neurons 0–5 — detect capital of Atlantis → write Paris; currency of Atlantis → Euro; and so on. Feed an address, one matmul, and out come Paris, Euro, Latin. No iteration, no demixing. To add a fact, add a row and a neuron — facts don't interfere, and capacity is in the number of neurons, not the demix budget. The underlying code is a NumPy array with the addresses stacked in and the facts inserted by hand.
Back to the opening demo: facts Gemma had never seen, written as a single value at the right address, read natively by the model's own forward pass — because the injection respected relations the model already knew. The trick is position: get it aligned to all the relevant relation variations so it fires where needed and doesn't interfere elsewhere. The model can read hundreds of such insertions; the danger is collision — once inserted facts clash, the model can no longer read them.
Which points at the real limitation: model dimensionality. The very packing mechanism that makes storage efficient is what bounds it. But — and this is the tease for a future video — you don't have to store the key–value pairs in the weights at all. As long as you respect the addressing sequence, the store can be external, and then it scales as far as you want.
"There is nothing more intelligent that's happening underneath the hood. This is how models do it."
"What the model is lacking isn't the storage — it's managing the storage fine. It's the demixing."
"It isn't actually reading a packed channel by peeling. It's reading each fact by an address."
"The model is packing to store but it's addressing to read... It never unpacks. It never decodes. It's just reading addresses."
"You don't need to store the key–value pairs within the weights. You can store them in an external store... as long as you respect the addressing sequence."
| Time | Topic |
|---|---|
| 00:00 | Intro |
| 02:30 | How packing works |
| 05:00 | Decoding |
| 09:30 | Polysemanticity |
| 11:00 | The fact layers |
| 14:30 | The compute ladder |
| 17:10 | Progressive addressing |
| 22:33 | The conveyor belt |
| 24:33 | Generalization |
| 27:00 | Hand-built FFN |
Named on screen in the video:
Credited in the video description only (“standing on the shoulders of”):