Back to Blog
4 min read
Serve compressed VLM inference from a container
"Build a container image with turboquant-vllm baked in, serve a vision-language model with 3.76x KV cache compression, and verify it works — in under five minutes."
The first turboquant-vllm release proved the algorithm works — pip install, one flag, 3.76x KV cache compression. But if you've ever set up a GPU inference environment from scratch, you know the real friction isn't the model or the framework. It's the CUDA toolkit version, the driver compatibility matrix, the pip packages that refuse to coexist.
v1.1.0 ships a Containerfile that eliminates that entire setup. Build the image once, and every run starts from a known-good state — vLLM, CUDA runtime, and the TQ4 compression plugin verified at build time.
This guide walks through building the container, serving a vision-language model with compressed inference, verifying it works, and optionally running it as a persistent systemd service.
- An NVIDIA GPU with drivers installed (tested on RTX 4090, 24 GB). AMD ROCm also works — adjust the device flag.
- Podman or Docker. Commands below use Podman. For Docker, swap
podmanfordocker. - Enough VRAM. Molmo2-8B needs ~24 GB at 6K context with
--gpu-memory-utilization 0.90. Molmo2-4B fits with longer contexts on the same card.
Clone the repo and build:
The Containerfile does two things: installs turboquant-vllm from PyPI into the official vLLM image, then verifies the plugin entry point registered correctly. If the entry point check fails, the build fails — you won't discover a misconfigured plugin at runtime.
The TURBOQUANT_VERSION build arg defaults to 1.1.0. Override it for future versions without touching the file.
One flag does all the work: --attention-backend CUSTOM. This tells vLLM to use the TQ4 backend instead of its default attention implementation. Everything else — model loading, tokenization, the OpenAI-compatible API — stays exactly the same.
The named volume (vllm-models) caches model weights between container restarts. Multi-gigabyte checkpoints download once.
Watch the container logs for the backend confirmation:
If you see FLASH_ATTN or XFORMERS instead, the plugin didn't register. Rebuild the image and check the entry point verification passed.
You can also confirm from inside a running container:
The container exposes the standard vLLM OpenAI-compatible API. Nothing changes on the client side:
Clients don't know — and don't need to know — that the KV cache is 3.76x compressed behind the API.
For production, Quadlet manages the container as a systemd service. Create ~/.config/containers/systemd/vllm-turboquant.container:
Reload and start:
The health check gives the model up to five minutes to load weights before marking the service unhealthy. Restart=always handles crashes and GPU driver hiccups automatically.
Build fails at entry point verification. The vLLM base image version may not match turboquant-vllm's requirements. Check PyPI for supported vLLM versions.
Container starts but uses FLASH_ATTN instead of CUSTOM. Confirm --attention-backend CUSTOM is in your run command. In Quadlet, it goes in the Exec= line after the model name.
OOM during prefill. TurboQuant compresses the KV cache, not model weights or activations. Peak memory during prefill is activation-dominated — compression savings appear during generation. Lower --max-model-len or use a smaller model variant.
You built a container image with turboquant-vllm baked in, served a vision-language model with 3.76x KV cache compression, verified the plugin was active, and optionally deployed it as a persistent systemd service — all from one Containerfile and one CLI flag.
The full API reference and additional usage guides are on the documentation site.