Back to Blog
TurboQuant on vision models
A series in 4 parts, published from to .
- Part 1 of 4explanation
6 min read
I ran TurboQuant on a vision model. The first output was garbage.
I implemented Google's TurboQuant algorithm for KV cache compression and validated it on Molmo2 video inference on an RTX 4090 — 3.76x compression with near-identical output at 1.78x overhead.
- Part 2 of 4explanation
5 min read
"Paper to PyPI in 72 hours: Building the first TurboQuant vLLM plugin"
"Google published TurboQuant at ICLR 2026 for text models. 72 hours later, turboquant-vllm was on PyPI — the first implementation validated on vision-language models and the first vLLM plugin. One flag to enable, 3.76x KV cache compression."
- Part 3 of 4how-to
4 min read
Serve compressed VLM inference from a container
"Build a container image with turboquant-vllm baked in, serve a vision-language model with 3.76x KV cache compression, and verify it works — in under five minutes."
- Part 4 of 4explanation
5 min read
"From one model to seven: Making TurboQuant model-portable"
"turboquant-vllm started as a Molmo2-only proof of concept. v1.3.0 validates seven model families — but getting there meant rewriting Triton kernels for non-standard head dimensions and teaching the cache about sliding window attention."