Chris Hay takes Qwen3.8-27B from roughly 55 GB down to an 18.8 GB GGUF running in llama.cpp on a laptop — then argues the label on that file is nearly meaningless. He didn’t quantize the model once; he made different decisions per component. “4-bit” is a code width, not a storage cost: the block lands around 4.5 bits per weight, and tensor-level overhead takes the decoder stack to 4.5012. Some surfaces can’t use the block format at all; others were held at 16 bits by choice, and some were dropped once measurement showed they hadn’t earned the extra bits. The conclusion: a model doesn’t have a precision, it has a precision program — and quantization is better understood as compilation.
vindex represent compiles an alternative representation of the same components into a new container — NVFP4 sitting beside BF16, not overwriting it. What you get is a catalog of alternatives, not yet a decision about what to run.q_proj” is quantizing by filename. The graph tells you what the tensor does.vindex verify (~13 min) checks the artifact; vindex export compiles it down to an ordinary 18.8 GB GGUF. After that the VINDEX3 file can disappear — llama.cpp knows nothing about the fidelity experiments, only the tensors, encodings, layouts and metadata they compiled into.The video opens on the finished thing: Qwen3.8-27B running under llama.cpp on a laptop. The original is about 55 GB; this one is under 20. Most of the language-model weights are using 4-bit codes. But, Hay argues, calling it a 4-bit model hides almost everything interesting about how the file was made — because he didn’t quantize it once. Some weights use four-bit codes, some don’t, some can’t, and some started higher until measurement showed they hadn’t earned the extra bits.
So the video isn’t about how to make a 4-bit number. It’s about a harder question: which information can this model afford to lose?
He walks down into the structure — 64 layers, the first layer’s feed-forward block (dense), and inside it the down projection: a matrix of weights, the actual thing to be quantized. The same tensor comes up identically in the browser view and in the terminal.
One detail is planted deliberately here. He asked for the down projection — the semantic component — not for a tensor whose name happens to contain “down.” That distinction pays off later.
Switching the representation to BF16 shows sixteen values with no error factor at all. Switching to the 4-bit projection shows the originals, the reconstructed values, and the gap between them.
Two things fall out. First, four bits is the code width, not the storage cost — the scales cost space too, putting the block at roughly 4.5 bits per weight, with tensor-level overhead taking the decoder stack to 4.5012. Second, it is lossy compression, not zip: distinct numbers can collapse onto the same value and there is no way back. Which reframes the whole exercise — the error isn’t the question, whether the model cares about it is.
Rather than trust an animation, he compiles a second representation of the same model. vindex represent takes the identical semantic components into a new container — about two minutes — and the comparison of BF16 against NVFP4 reproduces exactly what the visualisation showed.
The important part is that both columns live in the same container. Nothing was converted and nothing was replaced. The BF16 representation is still there; the NVFP4 representation sits beside it. What you have is a catalog of alternatives — and the choice of which one actually runs is made later, by the precision program, which binds each semantic component to a physical representation that exists.
The layer grid across the 64 layers shows the model isn’t even using the same token mixer in every layer. Layer 0’s mixer is a gated delta net; layer 3’s is gated attention, with query, key, value, output and output gate. The pattern is three delta layers then one gated softmax attention layer, repeated 16 times.
And the mixer isn’t a guess — the artifact describes it. vindex describe qwen-nvfp4 layer.0.mixer returns the same structure locally that the site showed, with the semantics and tensors attached; run it on layer 3 and the query/key/value of gated attention appear instead.
The delta layer does have queries, keys and values — they just aren’t the QKV of softmax attention. In the attention layer, the query has its own projection. In delta, the query lives inside a fused recurrent QKV projection.
Hence the callback to the opening. A quantization policy that says “find every tensor whose name contains q_proj” is matching on an implementation detail. The graph is what says what a tensor does — and if different parts of the model don’t even compute the same function, there’s no reason they should automatically get the same precision.
The compiled precision map for Qwen3.8-27B shows it directly. Inside the gated delta net, most weights are 4.5 bits — but not all. The delta convolution can’t use the block representation, so a few model surfaces stay at 16 bits. Set against a standard uniform NVFP4 map, where that convolution would be pushed to 4.5, the difference is an accuracy trade he chose not to make.
"So what is the precision of this model? It doesn't have one. It has a precision program, and this is its map."
And 16 bits doesn’t always mean the same thing. Sometimes it’s a structural constraint; sometimes it’s just a policy choice made on the current program. Whether those choices are good is a separate question — some were tested experimentally, and where extra precision didn’t buy enough behavior, he stopped spending the bits. The same settings are readable from the terminal via vindex precision, showing the 16-bit convolution running from layer 0 to 62.
A separate warning about how quantization gets evaluated. Running “the capital of France is” through the reference model returns Paris at 55%. Switching to a deliberately broken representation returns Paris again — but at 100%.
Same top token, and the distribution behind it completely annihilated. Matching the argmax is not evidence that a quantization is faithful.
Once a representation exists, vindex verify — about 13 minutes — reports whether the artifact is correct. Then vindex export lowers it out of the VINDEX3 format into a completely separate physical format, in this case GGUF: an 18.8 GB file, byte-identical to the one running at the top of the video. Loaded into llama.cpp, it answers the France question correctly, and at that point the VINDEX3 file can be thrown away.
llama.cpp doesn’t know about any of the fidelity experiments. It doesn’t know why the convolution stayed at 16 bits or why the recurrent controls ended up at four and a half. Those decisions were compiled away into tensors, encodings, layouts and metadata; the runtime just executes the result.
"Once you start asking that component by component, quantization stops looking like a compression setting and it starts to look like compilation."
The shape that leaves you with: one semantic model, multiple physical representations, a program choosing between them, and a runtime that executes what the program selected. Making 4-bit codes was the easy part; the hard question was what the model could afford to forget.
Hay picked these settings by hand. The tease for a future video is that he doesn’t have to — with representations and balancers, the right program can be searched for and optimized rather than chosen manually or applied uniformly. State a target — running on a Mac, 64 GB, 128 GB, speed prioritised over accuracy — pick or find the representation that suits it, and compile down to whichever runtime you want: llama.cpp, MLX, or LARQL natively.
The VINDEX3 specification, an ask feature for questions about the model, and a component/layer explorer are at vindex3.org; the explorer is currently tied to a demo model, with more to be added over time.
"Calling this a 4-bit model hides almost everything interesting about how that 19 GB file was created."
"So four bits is the code width, isn't the storage cost. The scales cost space too."
"So the error isn't the question. The question is: does the model care about the error in the first place?"
"We're not converting the model. We're not replacing the model. We're representing the same model differently."
"Tensor names are implementation details. The graph tells me what the tensor does."
"Just because it gives back the same token, it doesn't necessarily mean that a quantization is faithful."
"The difficult question was: what information could the model afford to forget?"
| Time | Topic |
|---|---|
| 00:00 | An 18.8 GB Qwen3.8-27B, and the question behind it |
| 00:50 | Opening the model: layer 0, the down projection |
| 01:52 | What “4-bit” actually costs — code width vs storage |
| 03:15 | vindex represent — one model, many representations |
| 04:56 | The stack grid: gated delta net vs gated attention |
| 06:38 | Why tensor names are the wrong handle |
| 07:30 | The precision map — a precision program, not a precision |
| 09:05 | Never judge a quantization by its top token |
| 10:15 | Verify, then export down to GGUF |
| 12:21 | Quantization as compilation |
| 13:12 | What’s next: searching for the right representation |
| 14:05 | vindex3.org walkthrough |
vindex represent — compile an alternative representation into the container (~2 min).vindex layers — list the per-layer mixer pattern.vindex describe <model> layer.N.mixer — describe a layer’s mixer, semantics and tensors.vindex precision — show the compiled precision settings per surface.vindex verify — check the artifact is correct (~13 min).vindex export — lower the representation into GGUF.q_proj; “GDN” → gated delta net; “po choice” → policy choice; “sematic” → semantic. At 02:44 the captions read “45012” — in context this is 4.5012 bits per weight, following the stated ~4.5-bit block cost. At 08:00 the captions read “what is the position of this model”; this is precision, following directly from the precision map. The exported file is stated as both “19 gig” and, specifically, 18.8 GB.