Back to Blog
Measuring model confidence
A series in 3 parts, published from to .
- Part 1 of 3explanation
11 min read
Jev looked underconfident on Banking77. Two of my labels were the reason.
I asked TypeSafe's Jev one yes-or-no question about a thousand public messages and checked whether its percentages matched what actually happened. On scam text messages they did. On bank support messages they looked far too cautious, until I read the messages themselves. The problem was two of my labels, not Jev.
- Part 2 of 3explanation
8 min read
My Gemma was 99.96 percent sure the total matched. It had not seen the receipt.
typevet is a small Python library, public today, that asks an open model typed questions and reads the answer as a probability. I asked Gemma 4 31B, on my own 24 GiB card, whether an expense claim matched a real receipt. Without the photo it said yes to every claim that showed a full number, almost certain every time. With the photo it caught all six wrong totals.
- Part 3 of 3explanation
6 min read
Gemma 4 read a love note as a scam. A rewritten question brought it within two messages of Jev.
I asked Jev and Gemma 4 31B the same scam question about 158 public text messages. Jev got 151 right. Gemma, on my own card, got 131, because it flagged love notes and everyday chat as fraud. Then I let an optimizer rewrite the question once for each model. Gemma climbed to 153, and on calibration its percentages ended closer to reality than Jev's.