Back to Blog
vramfit
A series in 6 parts, published from to .
- Part 1 of 6explanation
12 min read
I couldn't tell my quantized model from the baseline. The instruments could.
I shrank a 93 GB model onto a 24 GB card by measuring which layers survive being crushed, instead of guessing. Then I couldn't tell the result from the standard quant by talking to it — which turns out to be the whole point.
- Part 2 of 6explanation
8 min read
A different ceiling is a different recipe. I finally checked.
I claimed a smaller card gives you a different quantization recipe, not the same one squeezed, and then didn't prove it. Here's that claim run against the published price list — including the ceiling where the honest answer is that there's no dish.
- Part 3 of 6how-to
11 min read
Fit a model to the GPU you actually have
Measure a model's per-layer quantization damage, solve for a recipe that fits your card, and check the result before you trust it.
- Part 4 of 6explanation
13 min read
The 2-bit label was 4.5 bits inside. My 16 GiB card could tell.
The smallest 2-bit-labeled build of Nemotron 3.5 Lightning is 17.54 GiB — it doesn't fit a 16 GiB card, and its label names 12 of 417 tensors. I measured the model stack by stack instead, and got a 15.76 GiB pack that serves fully on-card at 16k context and beats the shelf's build on both damage metrics, while 1.78 GiB smaller.
- Part 5 of 6explanation
13 min read
Google's 4-bit Gemma already fit my 24 GiB card. I wanted the 20,000 tokens it left on the table.
The official Q4_0 build of Gemma 4 31B is 16.44 GiB and fits a 24 GiB card with room to spare, so "it fits" was never the claim. The budget is. I measured the decoder layer by layer, solved a 14.92 GiB pack that ties Google's build on four held-out benchmarks and wins one, and let the freed bytes buy context — 86,016 served tokens against 65,536, and 73,728 against 49,152 with an image aboard.
- Part 6 of 6how-to
11 min read
Serve Gemma 4 31B on a 24 GiB card with the context it was packed for
Download the published fit24gib pack, install llama.cpp at the build the numbers were measured on, and serve it at 86,016 tokens of text context or 73,728 with an image aboard. Every flag comes from the model card or the vramfit repo's own serve how-to, and the section at the end says what to do when your card does not reach the boundary.