Back to Blog

TurboQuant on vision models

A series in 4 parts, published from to .

  1. Part 1 of 4
    explanation

    6 min read

    I ran TurboQuant on a vision model. The first output was garbage.

    I implemented Google's TurboQuant algorithm for KV cache compression and validated it on Molmo2 video inference on an RTX 4090 — 3.76x compression with near-identical output at 1.78x overhead.

  2. Part 2 of 4
    explanation

    5 min read

    "Paper to PyPI in 72 hours: Building the first TurboQuant vLLM plugin"

    "Google published TurboQuant at ICLR 2026 for text models. 72 hours later, turboquant-vllm was on PyPI — the first implementation validated on vision-language models and the first vLLM plugin. One flag to enable, 3.76x KV cache compression."

  3. Part 3 of 4
    how-to

    4 min read

    Serve compressed VLM inference from a container

    "Build a container image with turboquant-vllm baked in, serve a vision-language model with 3.76x KV cache compression, and verify it works — in under five minutes."

  4. Part 4 of 4
    explanation

    5 min read

    "From one model to seven: Making TurboQuant model-portable"

    "turboquant-vllm started as a Molmo2-only proof of concept. v1.3.0 validates seven model families — but getting there meant rewriting Triton kernels for non-standard head dimensions and teaching the cache about sliding window attention."

© 2026 Alberto Nieto. All rights reserved.