Back to Blog
2026-09-29
8 min read
typevet is a small Python library, public today, that asks an open model typed questions and reads the answer as a probability. I asked Gemma 4 31B, on my own 24 GiB card, whether an expense claim matched a real receipt. Without the photo it said yes to every claim that showed a full number, almost certain every time. With the photo it caught all six wrong totals.
The claim said the receipt came to 646,329. The receipt says 664,329. Two digits swapped, the kind of slip anyone makes typing a number in a hurry.
I asked Gemma 4 31B, running on the graphics card under my desk, whether the claim matched the receipt. When I sent the claim without the photo, it said the total matched, 99.96 percent sure. It had never seen the receipt. When I attached the photo, it said the total did not match, 99.999 percent sure.
Both answers came out of typevet, a library I made public today.
In the Jev calibration post I described a model you do not talk to. You write the test it takes, like a fill-in-the-bubble sheet, and it tells you how likely each bubble is. Jev is TypeSafe's hosted model, built for exactly that job. It also reads nothing but text. TypeSafe's docs say so plainly: "Jev accepts text only. ... Images, audio, and video are not supported (yet)."
The questions I wanted to ask were about pictures. So I built typevet to ask the same kind of question of an open model that can see, running on my own card. Call it Jev at home. Jev at home can look at a receipt.
typevet has three question types, the same three shapes Jev answers:
typevet does not let the model write an answer and then try to parse it. Before the model writes anything, typevet reads how likely each allowed answer is as the very next word. The labels you did not offer get no vote. What comes back is a label with a probability for every option, or an error. It is never a paragraph you have to interpret.
Here is the question the receipt test asked, shortened where you see ....
It is data, not a prompt I wrote around the model:
The photo rides along with the text. The model sees both, and typevet reads the three probabilities.
I wanted a test I could check by eye. CORD is a public set of real shop receipts, photographed and labelled with their totals. I took six of them and wrote three made-up expense claims for each:
?, so no total can be read.That is 18 claims. I sent each one twice: once as text alone, and once with the receipt photo attached.
| Claim sent without the photo | Claim sent with the photo | |
|---|---|---|
| Right total, said match | 6 of 6 | 6 of 6 |
| Wrong total, said does not match | 0 of 6 | 6 of 6 |
| Smudged total, said not enough to tell | 6 of 6 | 6 of 6 |
Without the photo, the model said match to every claim that showed a full number, twelve of twelve, including all six wrong ones. On the wrong ones it was at least 99.96 percent sure. With the photo, it caught every swapped total, each at 99.998 percent or more.
It also refused to guess on the smudged claims in both runs. The question
tells it that a ? is a lost digit, and it listened.
A probability tells you how sure the model is. It does not tell you what the model looked at. Without the photo, that 99.96 percent rests on nothing but a tidy claim. With the photo, the 99.999 percent rests on the total printed on the receipt. A model that can see is the difference, and a text-only model cannot give you that.
typevet's tests make sure the photo is really there. In this run, attaching a receipt added between 228 and 1,108 prompt tokens, and the tests fail if the count does not grow.
Gemma 4 31B reads text and images. It is released under the Apache 2.0 licence, and it placed third among open models on LMArena when it launched. I chose it because it reads pictures, its licence lets me build on it, and it can be made to fit on one consumer card.
It fits on mine because of the 24 GiB pack I built for the last vramfit post. The file the receipt test loaded is byte for byte the one published on Hugging Face, served by llama.cpp on an RTX 4090. Nothing in the test left the machine.
I would rather not send a receipt to somebody else's server to find out whether two numbers agree.
For hosted work, typevet talks to vLLM. On one rented H100 running the full-precision Gemma 4 31B, the receipt test got all 18 claims right with the photo. Reversing the order of the three answer options changed none of the 18 answers.
For speed, I sent 480 public bank support messages through it, two questions each. One at a time, a message came back in 0.24 seconds at the median. With 64 in flight, the server answered 39.6 messages a second, with no errors.
I am not the only person who wanted typed answers about pictures. Jev launched on 2026-09-15, and open takes on the idea followed fast.
What typevet brings is a library, not a script. It runs the larger Gemma 4 31B on one 24 GiB card with llama.cpp, and the same code serves a rented H100 through vLLM. And its tests check that each image reached the model before any answer counts.
The receipt test is 18 made-up claims on six real receipts, one run on each server. It shows the photo reaching the model and deciding the answer. The probabilities are the model's own confidence, not a calibrated forecast. Checking calibration, as in the Jev post, takes hundreds of labelled examples.
Same model, same claim, the same near-certainty both times. The photo is what made it right.