<?xml version='1.0' encoding='utf-8'?>
<rss xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/" version="2.0">
  <channel>
    <title>alberto.codes</title>
    <link>https://alberto.codes/blog</link>
    <description>Thoughts on AI engineering, Python, career growth, and technical leadership.</description>
    <language>en-us</language>
    <lastBuildDate>Thu, 01 Oct 2026 00:00:00 GMT</lastBuildDate>
    <atom:link href="https://alberto.codes/feed.xml" rel="self" type="application/rss+xml" />
    <item>
      <title>Gemma 4 read a love note as a scam. A rewritten question brought it within two messages of Jev.</title>
      <link>https://alberto.codes/blog/2026-10-01-gemma-read-a-love-note-as-a-scam</link>
      <guid isPermaLink="true">https://alberto.codes/blog/2026-10-01-gemma-read-a-love-note-as-a-scam</guid>
      <pubDate>Thu, 01 Oct 2026 00:00:00 GMT</pubDate>
      <description>I asked Jev and Gemma 4 31B the same scam question about 158 public text messages. Jev got 151 right. Gemma, on my own card, got 131, because it flagged love notes and everyday chat as fraud. Then I let an optimizer rewrite the question once for each model. Gemma climbed to 153, and on calibration its percentages ended closer to reality than Jev's.</description>
      <content:encoded>&lt;p&gt;&amp;quot;I want to show you the world, princess :) how about europe?&amp;quot;&lt;/p&gt;
&lt;p&gt;That is a text message from a public scam dataset, and it is not a scam. I
asked Gemma 4 31B, running on the graphics card under my desk, whether it was
a scam, phishing or social-engineering attempt. It said yes, 99.98 percent
sure.&lt;/p&gt;
&lt;p&gt;Then I let an optimizer rewrite that one question. Asked the new way, Gemma
put the same message under 0.001 percent. Across 158 messages it had never
seen, it went from 131 right to 153, two short of Jev's 155.&lt;/p&gt;
&lt;h2&gt;The test&lt;/h2&gt;
&lt;p&gt;In the &lt;a href="https://alberto.codes/blog/2026-09-27-jev-looked-underconfident-two-labels-were-the-reason"&gt;Jev calibration post&lt;/a&gt;
I asked TypeSafe's Jev one yes-or-no question about text messages from
&lt;a href="https://huggingface.co/datasets/difraud/difraud"&gt;DIFrauD&lt;/a&gt;, a public set
labelled scam or not. In the &lt;a href="https://alberto.codes/blog/2026-09-29-gemma-was-sure-the-total-matched-it-had-not-seen-the-receipt"&gt;last post&lt;/a&gt;
I built typevet, which asks the same kind of question of Gemma 4 on my own
card.&lt;/p&gt;
&lt;p&gt;So I put Jev and the Jev at home side by side. Same question, same 158 messages, 36 of them
scams:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Is this message a scam, phishing or social-engineering attempt?&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;Jev got 151 right. Gemma got 131. It missed no scams at all, but it flagged
27 ordinary messages as scams. Love notes. &amp;quot;Got meh... When?&amp;quot; &amp;quot;I fetch yun
or u fetch?&amp;quot; Gemma read anything personal and a little cryptic as somebody
working an angle.&lt;/p&gt;
&lt;h2&gt;Letting each model rewrite its question&lt;/h2&gt;
&lt;p&gt;The question is one sentence. Small wording changes move these models a lot,
so I let an optimizer search for a better sentence, once for each model.&lt;/p&gt;
&lt;p&gt;The optimizer is &lt;a href="https://arxiv.org/abs/2507.19457"&gt;GEPA&lt;/a&gt;, run through
&lt;a href="https://github.com/Alberto-Codes/gepa-adk"&gt;gepa-adk&lt;/a&gt;, my package for it. It
works like an editor with a red pen. The model answers a few practice
messages. A second model, Qwen3.8-27B on the same desk, reads how those
answers scored and proposes a rewrite. The winning rewrite is the one that
scores best on 200 separate messages. Every rewrite had to fit in 94
characters, so it could not grow into an essay.&lt;/p&gt;
&lt;p&gt;Each model got its own run with the same settings and data. Gemma's took 34 minutes
on my card. Jev's took 27 minutes. Each run made 2,280 calls. The 158 messages in the
test were never shown to either run.&lt;/p&gt;
&lt;h2&gt;The two rewrites went opposite ways&lt;/h2&gt;
&lt;p&gt;&lt;img src="https://alberto.codes/typevet-gepa-rewrites.svg" alt="Two cards. Left, Gemma 4's rewrite: Label 1 only for obvious deceptive scam/phishing/social engineering; ignore personal text. Its example, the message I want to show you the world, princess, how about europe?, labelled not a scam, went from 99.98 percent scam to under 0.001 percent. Right, Jev's rewrite: Predict probability (0-1) that message is scam/phishing/unsolicited promo, offer, call, alert. Its example, a real-estate promotion ending For Best Deal Call, labelled a scam, went from 32 percent to 89 percent." /&gt;&lt;/p&gt;
&lt;p&gt;Gemma's rewrite narrowed the word &amp;quot;scam&amp;quot;. Obvious deception only, and ignore
personal text. With that wording, every love note in the test came back as
not a scam.&lt;/p&gt;
&lt;p&gt;Jev's rewrite widened it. It added unsolicited promotions, offers, calls and
alerts. DIFrauD counts a lot of spam as scam, and Jev's rewrite suggests its
optimizer picked that up. A real-estate promotion that ends &amp;quot;For Best Deal Call&amp;quot; is labelled a
scam in the dataset. Jev moved it from 32 percent to 89.&lt;/p&gt;
&lt;h2&gt;The result&lt;/h2&gt;
&lt;p&gt;&lt;img src="https://alberto.codes/typevet-gepa-results.svg" alt="Two charts, one row per judge, each showing the seed question as a hollow gray dot and the model's own rewrite as a filled green dot. Left, messages right out of 158: Jev 151 to 155, Gemma 4 on my RTX 4090 131 to 153, Gemma 4 on a rented H100 128 to 153. Right, average gap between the stated percentage and what happened, in points, lower is better: Jev 8.9 to 7.1, Gemma 4 on the 4090 17.0 to 3.4, Gemma 4 on the H100 18.5 to 3.4." /&gt;&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Judge&lt;/th&gt;
&lt;th&gt;Right, same question&lt;/th&gt;
&lt;th&gt;Right, own rewrite&lt;/th&gt;
&lt;th&gt;Average gap, same question&lt;/th&gt;
&lt;th&gt;Average gap, own rewrite&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Jev&lt;/td&gt;
&lt;td&gt;151 of 158&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;155&lt;/strong&gt; of 158&lt;/td&gt;
&lt;td&gt;8.9 points&lt;/td&gt;
&lt;td&gt;7.1 points&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Gemma 4 31B on my RTX 4090&lt;/td&gt;
&lt;td&gt;131 of 158&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;153&lt;/strong&gt; of 158&lt;/td&gt;
&lt;td&gt;17.0 points&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;3.4&lt;/strong&gt; points&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Gemma 4 31B on a rented H100&lt;/td&gt;
&lt;td&gt;128 of 158&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;153&lt;/strong&gt; of 158&lt;/td&gt;
&lt;td&gt;18.5 points&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;3.4&lt;/strong&gt; points&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;The average gap is the calibration measure from the Jev post: how far a
model's percentages sit from what actually happened. Zero is perfect.&lt;/p&gt;
&lt;p&gt;Jev still gets the most messages right. Gemma's percentages now mean more:
on that measure they sit closer to what actually happened than Jev's do,
an average gap of 3.4 points against 7.1. Jev keeps a small edge on the
other scores in the study, and every gap here is small at this size.&lt;/p&gt;
&lt;h2&gt;Learned on my card, worked on the big one&lt;/h2&gt;
&lt;p&gt;The rewrite was learned on the &lt;a href="https://alberto.codes/blog/2026-09-02-googles-4-bit-gemma-already-fit-my-card"&gt;24 GiB pack&lt;/a&gt;
on my RTX 4090, which stores most weights at 4 bits and some at 2 or 3. I
then ran it on Google's full-precision Gemma 4 31B on a rented H100.&lt;/p&gt;
&lt;p&gt;Same answer, scam or not, on all 158 messages. The same 153 right, the same
3.4-point gap. A wording found on a compressed model at home carried over to
the full model on a datacenter card without losing a single message.&lt;/p&gt;
&lt;p&gt;The H100 was also the fast one. A message came back in 0.17 seconds at the
median there, 1.1 seconds on my desk, and 0.1 seconds from Jev's hosted
service.&lt;/p&gt;
&lt;h2&gt;The five that are left&lt;/h2&gt;
&lt;p&gt;Two are promotions that DIFrauD calls scams and Gemma's narrower question now
lets through, like the real-estate one above. The other three are these,
labelled &lt;em&gt;not&lt;/em&gt; a scam in the dataset:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&amp;quot;Send me your id and password&amp;quot;&lt;/li&gt;
&lt;li&gt;&amp;quot;What's ur pin?&amp;quot;&lt;/li&gt;
&lt;li&gt;&amp;quot;Perhaps * is much easy give your account identification, so i will tomorrow at UNI&amp;quot;&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Gemma says scam to all three, and so would I. So did Jev with the original
question.&lt;/p&gt;
&lt;h2&gt;Scope&lt;/h2&gt;
&lt;p&gt;158 messages with 36 scams, one run for each judge. At this size one message
moves accuracy by more than half a point. Each rewrite is tuned to DIFrauD's
idea of a scam, so it describes this dataset, not scams in general.&lt;/p&gt;
&lt;h2&gt;Where it lives&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;The study:&lt;/strong&gt; &lt;a href="https://alberto-codes.github.io/typevet/explanation/gemma-and-jev-difraud/"&gt;Gemma 4 and Jev on DIFrauD&lt;/a&gt;,
with the settings, every number above and the per-message receipts.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;The optimizer:&lt;/strong&gt; &lt;a href="https://github.com/Alberto-Codes/gepa-adk"&gt;gepa-adk&lt;/a&gt;,
Apache-2.0 licence.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;The model:&lt;/strong&gt; &lt;a href="https://huggingface.co/Alberto-Codes/gemma-4-31B-it-fit24gib-GGUF"&gt;gemma-4-31B-it-fit24gib-GGUF&lt;/a&gt;,
the 24 GiB pack the rewrite was learned on.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;The release:&lt;/strong&gt; &lt;a href="https://github.com/Alberto-Codes/typevet/releases/tag/v0.5.0"&gt;typevet v0.5.0&lt;/a&gt;,
the version that shipped this study: the Jev judge option for wording
evolution, the held-out comparison harness, and native Gemma 4 framing for
text runs.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;One sentence, rewritten once on my own card, took Gemma from 131 to 153. The
princess can go to Europe.&lt;/p&gt;
</content:encoded>
    </item>
    <item>
      <title>My Gemma was 99.96 percent sure the total matched. It had not seen the receipt.</title>
      <link>https://alberto.codes/blog/2026-09-29-gemma-was-sure-the-total-matched-it-had-not-seen-the-receipt</link>
      <guid isPermaLink="true">https://alberto.codes/blog/2026-09-29-gemma-was-sure-the-total-matched-it-had-not-seen-the-receipt</guid>
      <pubDate>Tue, 29 Sep 2026 00:00:00 GMT</pubDate>
      <description>typevet is a small Python library, public today, that asks an open model typed questions and reads the answer as a probability. I asked Gemma 4 31B, on my own 24 GiB card, whether an expense claim matched a real receipt. Without the photo it said yes to every claim that showed a full number, almost certain every time. With the photo it caught all six wrong totals.</description>
      <content:encoded>&lt;p&gt;The claim said the receipt came to 646,329. The receipt says 664,329. Two
digits swapped, the kind of slip anyone makes typing a number in a hurry.&lt;/p&gt;
&lt;p&gt;I asked Gemma 4 31B, running on the graphics card under my desk, whether the
claim matched the receipt. When I sent the claim without the photo, it said
the total matched, 99.96 percent sure. It had never seen the receipt. When I
attached the photo, it said the total did not match, 99.999 percent sure.&lt;/p&gt;
&lt;p&gt;&lt;img src="https://alberto.codes/typevet-receipt-swap.svg" alt="A receipt photo cropped to its total line, which reads 664,329. Beside it, the claim: total 646329. Two answers from the same model below. Claim sent without the photo: matches, 99.96 percent. Claim sent with the photo: does not match, 99.999 percent." /&gt;&lt;/p&gt;
&lt;p&gt;Both answers came out of typevet, a library I made public today.&lt;/p&gt;
&lt;p&gt;In the &lt;a href="https://alberto.codes/blog/2026-09-27-jev-looked-underconfident-two-labels-were-the-reason"&gt;Jev calibration post&lt;/a&gt;
I described a model you do not talk to. You write the test it takes, like a
fill-in-the-bubble sheet, and it tells you how likely each bubble is. Jev is
TypeSafe's hosted model, built for exactly that job. It also reads nothing but
text. TypeSafe's &lt;a href="https://docs.typesafe.ai/concepts/state"&gt;docs&lt;/a&gt; say so
plainly: &amp;quot;Jev accepts text only. ... Images, audio, and video are not
supported (yet).&amp;quot;&lt;/p&gt;
&lt;p&gt;The questions I wanted to ask were about pictures. So I built typevet to ask
the same kind of question of an open model that can see, running on my own
card. Call it Jev at home. Jev at home can look at a receipt.&lt;/p&gt;
&lt;h2&gt;What typevet does&lt;/h2&gt;
&lt;p&gt;typevet has three question types, the same three shapes Jev answers:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Yes or no: is this message a scam?&lt;/li&gt;
&lt;li&gt;Pick one label: does the claim match the receipt, not match it, or is
there not enough to tell?&lt;/li&gt;
&lt;li&gt;Pick one level on a scale: how well does this answer follow the rubric?&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;typevet does not let the model write an answer and then try to parse it.
Before the model writes anything, typevet reads how likely each allowed answer
is as the very next word. The labels you did not offer get no vote. What
comes back is a label with a probability for every option, or an error. It
is never a paragraph you have to interpret.&lt;/p&gt;
&lt;p&gt;Here is the question the receipt test asked, shortened where you see &lt;code&gt;...&lt;/code&gt;.
It is data, not a prompt I wrote around the model:&lt;/p&gt;
&lt;pre&gt;&lt;code class="language-python"&gt;Choice(
    instructions=(
        &amp;quot;Look at the claimed amount in the expense claim before you look at &amp;quot;
        &amp;quot;the receipt. ... Compare with the receipt total only when the claim &amp;quot;
        &amp;quot;shows every digit of its amount. Otherwise choose insufficient_evidence.&amp;quot;
    ),
    criteria={
        &amp;quot;insufficient_evidence&amp;quot;: &amp;quot;The claimed amount is not fully readable ...&amp;quot;,
        &amp;quot;mismatch&amp;quot;: &amp;quot;The claim shows every digit of its amount and that amount differs from the receipt total.&amp;quot;,
        &amp;quot;match&amp;quot;: &amp;quot;The claim shows every digit of its amount and that amount equals the receipt total.&amp;quot;,
    },
)
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;The photo rides along with the text. The model sees both, and typevet reads
the three probabilities.&lt;/p&gt;
&lt;h2&gt;The receipt test&lt;/h2&gt;
&lt;p&gt;I wanted a test I could check by eye. &lt;a href="https://huggingface.co/datasets/naver-clova-ix/cord-v2"&gt;CORD&lt;/a&gt;
is a public set of real shop receipts, photographed and labelled with their
totals. I took six of them and wrote three made-up expense claims for each:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;one that states the receipt's real total,&lt;/li&gt;
&lt;li&gt;one that states a wrong total, with two digits swapped,&lt;/li&gt;
&lt;li&gt;one with digits smudged out, written as &lt;code&gt;?&lt;/code&gt;, so no total can be read.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;That is 18 claims. I sent each one twice: once as text alone, and once with
the receipt photo attached.&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;Claim sent without the photo&lt;/th&gt;
&lt;th&gt;Claim sent with the photo&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Right total, said match&lt;/td&gt;
&lt;td&gt;6 of 6&lt;/td&gt;
&lt;td&gt;6 of 6&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Wrong total, said does not match&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;0 of 6&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;6 of 6&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Smudged total, said not enough to tell&lt;/td&gt;
&lt;td&gt;6 of 6&lt;/td&gt;
&lt;td&gt;6 of 6&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;Without the photo, the model said match to every claim that showed a full
number, twelve of twelve, including all six wrong ones. On the wrong ones it
was at least 99.96 percent sure. With the photo, it caught every swapped total, each
at 99.998 percent or more.&lt;/p&gt;
&lt;p&gt;It also refused to guess on the smudged claims in both runs. The question
tells it that a &lt;code&gt;?&lt;/code&gt; is a lost digit, and it listened.&lt;/p&gt;
&lt;h2&gt;Why the photo is the whole point&lt;/h2&gt;
&lt;p&gt;A probability tells you how sure the model is. It does not tell you what the
model looked at. Without the photo, that 99.96 percent rests on nothing but a
tidy claim. With the photo, the 99.999 percent rests on the total printed on
the receipt. A model that can see is the difference, and a text-only model
cannot give you that.&lt;/p&gt;
&lt;p&gt;typevet's tests make sure the photo is really there. In this run, attaching a
receipt added between 228 and 1,108 prompt tokens, and the tests fail if the
count does not grow.&lt;/p&gt;
&lt;h2&gt;Why Gemma 4, and why on my own card&lt;/h2&gt;
&lt;p&gt;&lt;a href="https://huggingface.co/google/gemma-4-31B-it"&gt;Gemma 4 31B&lt;/a&gt; reads text and
images. It is released under the Apache 2.0 licence, and it placed third
among open models on &lt;a href="https://x.com/arena/status/2039739427715735645"&gt;LMArena&lt;/a&gt;
when it launched. I chose it because it reads pictures, its licence lets me build on it, and it can be
made to fit on one consumer card.&lt;/p&gt;
&lt;p&gt;It fits on mine because of the
&lt;a href="https://alberto.codes/blog/2026-09-02-googles-4-bit-gemma-already-fit-my-card"&gt;24 GiB pack I built for the last vramfit post&lt;/a&gt;.
The file the receipt test loaded is byte for byte the one published on
&lt;a href="https://huggingface.co/Alberto-Codes/gemma-4-31B-it-fit24gib-GGUF"&gt;Hugging Face&lt;/a&gt;,
served by llama.cpp on an RTX 4090. Nothing in the test left the machine.&lt;/p&gt;
&lt;p&gt;I would rather not send a receipt to somebody else's server to find out
whether two numbers agree.&lt;/p&gt;
&lt;h2&gt;The same questions on a rented H100&lt;/h2&gt;
&lt;p&gt;For hosted work, typevet talks to vLLM. On one rented H100 running the
full-precision Gemma 4 31B, the receipt test got all 18 claims right with the
photo. Reversing the order of the three answer options changed
none of the 18 answers.&lt;/p&gt;
&lt;p&gt;For speed, I sent 480 public bank support messages through it, two questions
each. One at a time, a message came back in 0.24 seconds at the median. With
64 in flight, the server answered 39.6 messages a second, with no errors.&lt;/p&gt;
&lt;h2&gt;Who else does this&lt;/h2&gt;
&lt;p&gt;I am not the only person who wanted typed answers about pictures. Jev
launched on 2026-09-15, and open takes on the idea followed fast.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href="https://github.com/TypeLLM/TypeLLM"&gt;TypeLLM&lt;/a&gt; is where typevet's text
side comes from: how a JSON Schema becomes a set of decisions, and how each
decision is scored from the next-word probabilities. TypeLLM added image
input on 2026-09-24, on SGLang, tested with Qwen3.8-27B. It returns
probabilities for yes-or-no and pick-one fields, and its README example
reads a receipt too. typevet's image path was built separately and shares
no code with it. I worked it out on llama.cpp's own request format, for
Gemma 3 first and Gemma 4 the same day.&lt;/li&gt;
&lt;li&gt;&lt;a href="https://allanrbo.blogspot.com/2026/09/a-jev-like-wrapper-for-llms-including.html"&gt;A Jev-like wrapper for LLMs, including vision models&lt;/a&gt;,
a blog post from 2026-09-25, rebuilds Jev's three question types over
webcam images in one script, with Gemma 4 12B on llama.cpp and a 24 GiB
card.&lt;/li&gt;
&lt;li&gt;&lt;a href="https://linzhiqiu.github.io/papers/vqascore/"&gt;VQAScore&lt;/a&gt; has read the
probability of &amp;quot;Yes&amp;quot; from a vision model since 2024, to score whether an
image matches a caption.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;What typevet brings is a library, not a script. It runs the larger Gemma 4
31B on one 24 GiB card with llama.cpp, and the same code serves a rented H100
through vLLM. And its tests check that each image reached the model before
any answer counts.&lt;/p&gt;
&lt;h2&gt;Scope&lt;/h2&gt;
&lt;p&gt;The receipt test is 18 made-up claims on six real receipts, one run on each
server. It shows the photo reaching the model and deciding the answer. The
probabilities are the model's own confidence, not a calibrated forecast.
Checking calibration, as in the Jev post, takes hundreds of labelled
examples.&lt;/p&gt;
&lt;h2&gt;Where it lives&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;The code:&lt;/strong&gt; &lt;a href="https://github.com/Alberto-Codes/typevet"&gt;typevet on GitHub&lt;/a&gt;,
MIT licence.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;The docs:&lt;/strong&gt; &lt;a href="https://alberto-codes.github.io/typevet/"&gt;alberto-codes.github.io/typevet&lt;/a&gt;,
including how the image reaches each server and the public receipt for
the receipt test's pass.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;The model:&lt;/strong&gt; &lt;a href="https://huggingface.co/Alberto-Codes/gemma-4-31B-it-fit24gib-GGUF"&gt;gemma-4-31B-it-fit24gib-GGUF&lt;/a&gt;,
the 24 GiB pack the receipt test ran on.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;The release:&lt;/strong&gt; &lt;a href="https://github.com/Alberto-Codes/typevet/releases/tag/v0.1.0"&gt;typevet v0.1.0&lt;/a&gt;,
the first public version: yes-or-no, pick-one-label and scale questions over
text and images, on llama.cpp and vLLM, plus JSON output checked against a
schema you supply.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Same model, same claim, the same near-certainty both times. The photo is what
made it right.&lt;/p&gt;
</content:encoded>
    </item>
    <item>
      <title>Jev looked underconfident on Banking77. Two of my labels were the reason.</title>
      <link>https://alberto.codes/blog/2026-09-27-jev-looked-underconfident-two-labels-were-the-reason</link>
      <guid isPermaLink="true">https://alberto.codes/blog/2026-09-27-jev-looked-underconfident-two-labels-were-the-reason</guid>
      <pubDate>Sun, 27 Sep 2026 00:00:00 GMT</pubDate>
      <description>I asked TypeSafe's Jev one yes-or-no question about a thousand public messages and checked whether its percentages matched what actually happened. On scam text messages they did. On bank support messages they looked far too cautious, until I read the messages themselves. The problem was two of my labels, not Jev.</description>
      <content:encoded>&lt;p&gt;When a weather forecast says 30 percent chance of rain, it should rain on
about three of every ten days it says that. If it rains on nine of those ten
days, the forecast is too cautious. If it rains on none, it is too alarmed.
Forecasters call this being &lt;em&gt;calibrated&lt;/em&gt;: the percentages mean what they say.&lt;/p&gt;
&lt;p&gt;&lt;img src="https://alberto.codes/jev-forecast.svg" alt="Three rows of ten days, each day forecast at a 30 percent chance of rain. Top row, calibrated: it rained on 3 of the 10 days. Middle row, too cautious: it rained on 9 of the 10. Bottom row, too alarmed: it rained on none." /&gt;&lt;/p&gt;
&lt;p&gt;I wanted to know whether that holds for Jev, a model from TypeSafe that
answers yes-or-no questions with a percentage instead of a yes or a no. So I
asked it one question about a thousand public messages and compared its
percentages with what the messages actually were.&lt;/p&gt;
&lt;p&gt;On one set of messages, it looked badly miscalibrated. Where Jev said 20 to
50 percent, the real answer turned out to be yes almost nine times in ten.
That looked like a forecaster saying &amp;quot;probably not&amp;quot; before a week of rain.
Then I read the messages, and the mistake was mine.&lt;/p&gt;
&lt;h2&gt;What Jev says it does&lt;/h2&gt;
&lt;p&gt;Jev does not chat, write or explain itself. You give it some text and a
question, and it returns a number. For a yes-or-no question, TypeSafe's
&lt;a href="https://docs.typesafe.ai/primitives/noul"&gt;docs&lt;/a&gt; define that number as the
probability that the answer is yes.&lt;/p&gt;
&lt;p&gt;When I explain it to people, I say you don't talk to Jev. You write the test
it takes. Think of the fill-in-the-bubble exams from school, the Iowa tests or
an SAT answer sheet. Your job is to write the question and the answer choices.
Jev fills in the bubbles, except that instead of filling in one, it tells you
how likely each choice is. A yes-or-no question is a test with two bubbles.
With Jev, the work is in writing the question, not in a conversation.&lt;/p&gt;
&lt;p&gt;&lt;img src="https://alberto.codes/jev-bubble-test.svg" alt="Two cards. Left, labelled you write: the text, I have multiple charges on the same transaction; the question, Does the customer report a transaction they did not authorize?; and two empty bubbles, Yes and No. An arrow leads to the right card, labelled Jev fills in: the same bubbles shaded by how likely each is, Yes 16 percent and No 84 percent." /&gt;&lt;/p&gt;
&lt;p&gt;TypeSafe says its models are
&lt;a href="https://docs.typesafe.ai/concepts/system-one"&gt;&amp;quot;trained for calibrated decisions&amp;quot;&lt;/a&gt;,
and it is careful about what that promises: &amp;quot;Calibration is measured across
groups of predictions; it does not guarantee that an individual answer is
correct.&amp;quot; One message scored at 30 percent can be yes or no. But across a
large group of messages scored near 30 percent, about 30 percent should be
yes. A group is something I can check.&lt;/p&gt;
&lt;h2&gt;The test&lt;/h2&gt;
&lt;p&gt;I used two free public datasets where people had already labelled every
message.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;DIFrauD&lt;/strong&gt;: text messages, each labelled scam or not. I asked Jev: &amp;quot;Is this
message a scam, phishing or social-engineering attempt?&amp;quot;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Banking77&lt;/strong&gt;: short messages people send to a bank's support desk, sorted
into 77 topics such as &amp;quot;card arrived&amp;quot; or &amp;quot;exchange rate&amp;quot;. There is no scam
label here, so I made one. I picked six topics that sounded like &amp;quot;someone
took money I did not agree to&amp;quot; and counted them as yes. I asked Jev: &amp;quot;Does
the customer report a transaction they did not authorize?&amp;quot;&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Then I sorted Jev's answers into ten buckets: 0 to 10 percent, 10 to 20, and
so on. For each bucket I compared the average percentage Jev gave with how
often the answer was really yes. The average gap across all buckets is the
&lt;em&gt;calibration error&lt;/em&gt;. Zero would be perfect. I also measured something simpler:
whether Jev at least scored the yes messages higher than the no messages. Call
that the &lt;em&gt;sorting score&lt;/em&gt;. 1.0 means every yes message scored above every no
message.&lt;/p&gt;
&lt;h2&gt;The scam texts: close to the forecast&lt;/h2&gt;
&lt;p&gt;On the 500 text messages, Jev behaved like a good forecaster. The average gap
was 7 percentage points, and the sorting score was 0.995.&lt;/p&gt;
&lt;p&gt;The extremes were clean. Jev put 369 messages below 20 percent, and not one
of them was a scam. It put 47 messages at 90 percent or above, and 46 of them
were.&lt;/p&gt;
&lt;p&gt;&lt;img src="https://alberto.codes/jev-reliability.svg" alt="Three charts side by side. Each plots Jev's average percentage in ten buckets against how often the answer was really yes, with a dashed diagonal line where the two would match. Left: scam text messages, 500 messages, average gap 7 points, dots close to the line. Middle: bank messages with my six topics counted as yes, 480 messages, average gap 17 points; between 20 and 50 percent predicted, 87.5 to 93 percent were labelled yes, far above the line. Right: the same 480 answers with two topics recounted as no, average gap 6 points, dots back on the line." /&gt;&lt;/p&gt;
&lt;p&gt;In the chart, Jev's percentage runs along the bottom and the real share of
yes answers runs up the side. A dot on the dashed line means Jev's percentage matched reality
for that bucket. A dot above the line means more yeses than Jev predicted.
The left panel stays close to the line. The small wobbles in the middle come
from small buckets: a few buckets hold only 7 or 8 messages, so one message
moves the result by more than 10 points.&lt;/p&gt;
&lt;h2&gt;The bank messages: a forecaster saying &amp;quot;probably not&amp;quot; before the rain&lt;/h2&gt;
&lt;p&gt;On the 480 bank messages, the sorting score was still excellent at 0.986. Jev
put the yes messages above the no messages almost every time. But the
percentages themselves were off, with an average gap of 17 points, and always
in the same direction. None of the no messages scored 50 percent or more.
But 93 of the 240 yes messages scored below 50 percent. Jev looked too
cautious, as the middle panel shows.&lt;/p&gt;
&lt;p&gt;Splitting the yes messages by topic showed where the gap came from. Four
topics matched what I had asked about. Two did not. Each topic had 40
messages:&lt;/p&gt;
&lt;p&gt;&lt;img src="https://alberto.codes/jev-topics.svg" alt="Horizontal bars, one per topic of 40 messages, with a dashed line at 50 percent. A card payment I don't recognise, 83 percent, 36 of 40 at 50 percent or more. A direct debit I don't recognise, 82 percent, 35 of 40. A cash withdrawal I don't recognise, 80 percent, 34 of 40. My card may be compromised, 71 percent, 30 of 40. An extra charge on my statement, 36 percent, 8 of 40. A transaction charged twice, 29 percent, 4 of 40. The first four sit well above 50 percent; the last two sit below it." /&gt;&lt;/p&gt;
&lt;h2&gt;Reading the messages&lt;/h2&gt;
&lt;p&gt;Here are six messages from those two topics. My answer key said yes to every
one. Jev's percentage is beside each:&lt;/p&gt;
&lt;p&gt;&lt;img src="https://alberto.codes/jev-answer-key.svg" alt="Six messages from the two topics I counted as yes, each with Jev's percentage and my answer key, which says yes to all six. Charged twice: I have multiple charges on the same transaction, 16 percent. What are my remedies if I think I was charged twice for the same expense, 16 percent. Extra charge: When will the $1 transaction be credited to me, 7 percent. Why are there so many fees on my statement, 8 percent. There is a fee I don't recognize on my statement, 78 percent. I do not remember purchasing anything for 1 pound, and it is on my statement, 87 percent." /&gt;&lt;/p&gt;
&lt;p&gt;Someone who was charged twice agreed to the purchase. They are reporting a
billing mistake, not a transaction they never made. Jev said &amp;quot;probably not&amp;quot;,
and for the question I asked, that is the right answer. My label said yes.&lt;/p&gt;
&lt;p&gt;The &amp;quot;extra charge&amp;quot; topic is more mixed, and Jev kept up with it message by
message. The two about a fee or purchase the customer does not recognise
scored high. The two questions about a refund and about fees scored low. My
label treated all 40 the same because they share a topic name.&lt;/p&gt;
&lt;p&gt;In test terms, Jev answered the question printed on the page. The answer key I
graded it with was written for a slightly different question.&lt;/p&gt;
&lt;h2&gt;The same answers, counted fairly&lt;/h2&gt;
&lt;p&gt;I kept every one of Jev's 480 answers and changed only my labels, counting
those two topics as no. The average gap fell from 17 points to 6, which is
better than on the scam texts. That is the right-hand panel of the chart.&lt;/p&gt;
&lt;p&gt;I owe you one caveat about that 6. I chose to relabel those two topics after
seeing the results, and I checked the fix on the same answers. It is not a
fair new score for Jev. What it does show is that most of the 17-point gap
came from my labels answering a different question from the one I asked Jev.&lt;/p&gt;
&lt;p&gt;A third topic is borderline too. A compromised card does not always mean a
transaction the customer did not make, and Jev put 10 of those 40 messages
below 50 percent. I left that topic as yes. If I kept moving topics until the
number looked good, I would be tuning my labels to flatter the model.&lt;/p&gt;
&lt;h2&gt;What the two scores told me&lt;/h2&gt;
&lt;p&gt;The only thing I changed was two topics' labels. Here is what that did to
each score:&lt;/p&gt;
&lt;p&gt;&lt;img src="https://alberto.codes/jev-two-scores.svg" alt="Two small charts. Left, the sorting score, where higher is better: 0.986 with my first labels and 0.972 after relabelling two topics, nearly the same. Right, the average gap, where lower is better: 17 points with my first labels and 6 points after relabelling, about a third. Changing two topics' labels barely moved the sorting score and cut the average gap by about two thirds." /&gt;&lt;/p&gt;
&lt;p&gt;So the first result was not a measurement of Jev on its own. It measured
Jev's answers against my idea of what the answers should be. On the scam
texts, the labels came from the dataset's authors, answering nearly the same
question I asked. On the bank messages they came from me, answering a
slightly different question: which topics involve a charge the customer
questions.&lt;/p&gt;
&lt;h2&gt;What this does not show&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;This is two groups of about 500 messages, one run each. Several buckets
hold fewer than 20 messages, which is small.&lt;/li&gt;
&lt;li&gt;The bank sample is half yes by design. In the full dataset, those six topics
are 7.8 percent of messages, 240 of 3,080. A more natural mix would give a
different number.&lt;/li&gt;
&lt;li&gt;DIFrauD has several kinds of text. I used only the text messages.&lt;/li&gt;
&lt;li&gt;The same messages sent three days apart gave nearly the same results: an
average gap of 17 points both times on the bank messages, and 7 both times
on the scam texts. A handful of messages moved between buckets.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2&gt;For the curious&lt;/h2&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;What&lt;/th&gt;
&lt;th&gt;Detail&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Model&lt;/td&gt;
&lt;td&gt;TypeSafe &lt;code&gt;jev-1.13.0&lt;/code&gt;, called on 2026-09-27&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Client&lt;/td&gt;
&lt;td&gt;&lt;a href="https://github.com/Alberto-Codes/judgevet"&gt;judgevet&lt;/a&gt; 0.13.0, my Python client for Jev&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Calls&lt;/td&gt;
&lt;td&gt;980, all successful&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Scam texts&lt;/td&gt;
&lt;td&gt;&lt;a href="https://huggingface.co/datasets/difraud/difraud"&gt;DIFrauD&lt;/a&gt;, MIT licence: 500 messages from the SMS test split, 98 of them scams&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Bank messages&lt;/td&gt;
&lt;td&gt;&lt;a href="https://huggingface.co/datasets/PolyAI/banking77"&gt;Banking77&lt;/a&gt;, CC BY 4.0 licence: all 240 test messages from my six topics, plus 240 others chosen with a fixed seed&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;My six topics&lt;/td&gt;
&lt;td&gt;&lt;code&gt;card_payment_not_recognised&lt;/code&gt;, &lt;code&gt;direct_debit_payment_not_recognised&lt;/code&gt;, &lt;code&gt;cash_withdrawal_not_recognised&lt;/code&gt;, &lt;code&gt;compromised_card&lt;/code&gt;, &lt;code&gt;extra_charge_on_statement&lt;/code&gt;, &lt;code&gt;transaction_charged_twice&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Average gap&lt;/td&gt;
&lt;td&gt;Expected calibration error (ECE), ten equal-width buckets weighted by size: 0.073 scam texts; 0.174 bank messages as first labelled; 0.056 relabelled&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Sorting score&lt;/td&gt;
&lt;td&gt;AUROC: 0.995 scam texts; 0.986 bank messages as first labelled; 0.972 relabelled&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;TypeSafe calls a yes-or-no question a Noul, so judgevet does too, and the
answer's &lt;code&gt;.noul&lt;/code&gt; field is Jev's percentage as a number between 0 and 1. One
question is a few lines:&lt;/p&gt;
&lt;pre&gt;&lt;code class="language-python"&gt;import os

from judgevet import HTTPSystemOneAdapter, Noul

question = Noul(instructions=&amp;quot;Does the customer report a transaction they did not authorize?&amp;quot;)

with HTTPSystemOneAdapter(api_key=os.environ[&amp;quot;JEV_API__KEY&amp;quot;]) as client:
    response = client.system_one(
        state=&amp;quot;I have multiple charges on the same transaction.&amp;quot;,
        questions={&amp;quot;unauthorized&amp;quot;: question},
    )

print(response.nouls[&amp;quot;unauthorized&amp;quot;].noul)
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;The 17-point gap came apart when I stopped looking at the chart and read ten
messages from the buckets where it was widest.&lt;/p&gt;
</content:encoded>
    </item>
    <item>
      <title>The scan lost one letter, and a derivative was crowned the mother sauce</title>
      <link>https://alberto.codes/blog/2026-09-08-the-scan-lost-one-letter-and-crowned-a-derivative</link>
      <guid isPermaLink="true">https://alberto.codes/blog/2026-09-08-the-scan-lost-one-letter-and-crowned-a-derivative</guid>
      <pubDate>Tue, 08 Sep 2026 00:00:00 GMT</pubDate>
      <description>The 1907 scan reads entry 22 as BROWN SAUCE OR ESPAQNOLE, so no name in that witness held the word Espagnole except a derivative's, and saucier crowned LENTEN ESPAGNOLE as a mother with twelve preparations beneath it. The ticket blamed ranking. Measuring that fix is what found the real defect: the lookup checked that a base comes before its derivatives without ever checking that the base was there.</description>
      <content:encoded>&lt;p&gt;Here is what &lt;code&gt;saucier tree&lt;/code&gt; printed for the Espagnole family in the 1907
scan, before this week:&lt;/p&gt;
&lt;pre&gt;&lt;code class="language-console"&gt;$ uv run saucier tree espagnole --source escoffier-1907
LENTEN ESPAGNOLE  [espagnole]
├── HALF GLAZE  (en)
│   ├── SAUCE BORDELAISE  (fr)
│   ├── BROWN CHAUD=FROID SAUCE  (en)
│   ├── DEVILLED SAUCE  (en)
│   ├── ITALIAN SAUCE  (en)
│   ├── LYONNAISE SAUCE  (en)
│   ├── MADEIRA SAUCE  (en)
│   ├── PERIQUEUX SAUCE  (en)
│   ├── PIQUANTE SAUCE  (en)
│   └── ROBERT SAUCE  (en)
├── ORDINARY POIVRADE SAUCE  (en)
└── POIVRADE SAUCE FOR VENISON  (en)
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Twelve preparations, and the sauce at the head of them is a Lenten
variation — the fish-day Espagnole, made when the ordinary one will not do.
Escoffier's entry for it opens by doubting whether it needs to exist at all.
It is a derivative wearing the crown of the base it derives from.&lt;/p&gt;
&lt;h2&gt;One letter&lt;/h2&gt;
&lt;p&gt;The heading it should have taken:&lt;/p&gt;
&lt;pre&gt;&lt;code class="language-console"&gt;$ sed -n 1730p corpus/escoffier-1907.txt
22— BROWN  SAUCE  OR  ESPAQNOLE
$ sed -n 1392p corpus/escoffier-1909.txt
22—BROWN SAUCE OR ESPAGNOLE
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;&lt;code&gt;ESPAQNOLE&lt;/code&gt;. The scanner read the G as a Q, which it does all through this
witness: &lt;code&gt;16— POULTRY QLAZE&lt;/code&gt; at line 1539, &lt;code&gt;38— QENEVOISE SAUCE&lt;/code&gt; at 2150,
&lt;code&gt;39— QRAND-VENEUR SAUCE&lt;/code&gt; at 2222, &lt;code&gt;46— PIQNONS SAUCE&lt;/code&gt; at 2300. Three of
those have been sitting in the &lt;code&gt;diff&lt;/code&gt; output since the scan was added, as
rows labelled &lt;code&gt;ocr-suspected&lt;/code&gt;, which is exactly what they are:&lt;/p&gt;
&lt;pre&gt;&lt;code class="language-console"&gt;  ocr-suspected   piqnons-sauce ~ pignons-sauce           PIQNONS SAUCE / PIGNONS SAUCE
  ocr-suspected   qenevoise-sauce ~ genevoise-sauce       QENEVOISE SAUCE / GENEVOISE SAUCE
  ocr-suspected   qrand-veneur-sauce ~ grand-veneur-sauce QRAND-VENEUR SAUCE / GRAND-VENEUR SAUCE
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Harmless, and useful: that label is what
&lt;a href="https://alberto.codes/blog/2026-09-01-i-added-a-second-copy-of-the-same-book"&gt;the second-copy post&lt;/a&gt;
added the second witness to produce. &lt;code&gt;ESPAQNOLE&lt;/code&gt; is not on that list. It is
not anywhere in the diff, because entry 22's heading gives it a second name
and both witnesses key it on that one:&lt;/p&gt;
&lt;pre&gt;&lt;code class="language-console"&gt;$ uv run saucier show espaqnole --source escoffier-1907
BROWN SAUCE OR ESPAQNOLE
entry 22, line 1730, ocr of escoffier-1907
  term  BROWN SAUCE  [en]  brown-sauce
  term  ESPAQNOLE  [en]  espaqnole
  parent  brown-roux
  procedure  (unrecorded)
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;The base is there, one letter away, catalogued as &lt;code&gt;brown-sauce&lt;/code&gt; in both
printings and carrying the same &lt;code&gt;brown-roux&lt;/code&gt; parent in both, which is why
the diff has nothing to say about it. What entry 22 does not have is the word &lt;code&gt;espagnole&lt;/code&gt;
anywhere among its names. So when the catalogue was asked which preparation
the mother concept &lt;code&gt;espagnole&lt;/code&gt; names, every exact match failed, and the
search fell back to the rule for partial matches: the concept has to appear
as a whole run of words inside a name. In the 1907 witness exactly one name
satisfies that.&lt;/p&gt;
&lt;pre&gt;&lt;code class="language-text"&gt;LENTEN ESPAGNOLE   entry 24, line 1795
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;&lt;img src="https://alberto.codes/saucier-espaqnole-mother.svg" alt="Two states of the same lookup. Top, the heading of entry 22 as each witness carries it: the 1907 scan reads 22 em-dash BROWN SAUCE OR ESPAQNOLE at line 1730, with the Q drawn in orange, and the 1909 transcription reads ESPAGNOLE at line 1392. Below, two columns. Left, before: the only candidate for the mother espagnole is LENTEN ESPAGNOLE, entry 24 at line 1795, crowned because it is the only run match and there is nothing to rank, and the tree heads on LENTEN ESPAGNOLE with HALF GLAZE and its nine derivatives and the two Poivrade sauces beneath it. Right, after: that candidate is struck out because its opening paragraph states Espagnole, no candidate is left, the mother stays uncatalogued, and the tree heads on the bare concept espagnole with LENTEN ESPAGNOLE demoted to a child alongside the others. Caption: the guard removed one candidate in this witness, the 1907 census moved from 140 / 50 / 90 to 140 / 51 / 89, the 1909 catalogue came out byte-identical, and the heading was never repaired." /&gt;&lt;/p&gt;
&lt;h2&gt;The rule had a premise it never checked&lt;/h2&gt;
&lt;p&gt;&lt;a href="https://github.com/Alberto-Codes/saucier/blob/1664e75e95c6806ce604cbab0436e4f1bdaab1e0/docs/adr/0008-a-parent-may-be-any-catalogued-preparation.md"&gt;ADR-0008&lt;/a&gt;
says a mother binds to the first preparation, in source order, that answers
to its name. The reasoning is a fact about the book: Escoffier presents a
base before its derivatives, so among several names carrying the word, the
earliest one is the base. That is true of this book, and it is the reason
&lt;code&gt;veloute&lt;/code&gt; reaches &lt;code&gt;ORDINARY VELOUTÉ SAUCE&lt;/code&gt; at entry 25 rather than
&lt;code&gt;ALLEMANDE SAUCE OR THICKENED VELOUTÉ&lt;/code&gt; at entry 27.&lt;/p&gt;
&lt;p&gt;The premise underneath it is that the base is among the candidates at all.
The lookup sorted the hits by where they appear and returned the first one.
It never asked whether the list it was sorting contained the thing it was
looking for. When the scanner takes the base's only copy of the name, the
list still has entries in it, the ordering rule still fires, and the code
still returns something — with no less confidence than when it is right.&lt;/p&gt;
&lt;p&gt;There is a second casualty in the same failure. Entry 24 opens like this:&lt;/p&gt;
&lt;pre&gt;&lt;code class="language-console"&gt;$ sed -n 1797,1798p corpus/escoffier-1907.txt
Practical  men  are  not  agreed  as  to  the  need  of  Lenten
Espagnole.  The  ordinary  Espagnole  being  really  a  neutral
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;The first sentence states Espagnole on its own, because the run sits inside
&lt;code&gt;Lenten Espagnole&lt;/code&gt;, and the sentence after it names the ordinary Espagnole
outright. The parent resolver reads exactly that kind of sentence and would
have recorded &lt;code&gt;espagnole&lt;/code&gt; as this entry's parent. It did not, because it discards a statement that names
the entry doing the stating, and the mother lookup was telling it that
&lt;code&gt;espagnole&lt;/code&gt; &lt;em&gt;was&lt;/em&gt; entry 24. So the same defect that gave Lenten Espagnole a
crown also cost it the parent it plainly states. &lt;code&gt;saucier show&lt;/code&gt; reported
&lt;code&gt;parent (unresolved)&lt;/code&gt; and &lt;code&gt;stated no candidate&lt;/code&gt; for an entry whose first
sentence names its base.&lt;/p&gt;
&lt;h2&gt;The proposed fix was aimed at ranking&lt;/h2&gt;
&lt;p&gt;The ticket that opened this proposed that a mother bind only on an exact
name match. If nothing in the witness is named exactly &lt;code&gt;espagnole&lt;/code&gt;, bind
nothing, and the bad crown disappears.&lt;/p&gt;
&lt;p&gt;It does. There are ten mother bindings across the two witnesses: five
mothers — &lt;code&gt;bechamel&lt;/code&gt;, &lt;code&gt;espagnole&lt;/code&gt;, &lt;code&gt;hollandaise&lt;/code&gt;, &lt;code&gt;tomato&lt;/code&gt;, &lt;code&gt;veloute&lt;/code&gt; —
resolved once per witness. Exactly one of the ten is an exact name match,
the 1909 &lt;code&gt;ESPAGNOLE&lt;/code&gt; that Escoffier prints as the second term of entry 22's
heading. The other nine reach their base through the run rule, because what
the book prints is &lt;code&gt;BÉCHAMEL SAUCE&lt;/code&gt;, &lt;code&gt;HOLLANDAISE SAUCE&lt;/code&gt;, &lt;code&gt;TOMATO SAUCE&lt;/code&gt;.
Binding on exact names only removes nine bindings in order to remove one
wrong one.&lt;/p&gt;
&lt;p&gt;The census will not tell you that. Run the proposal and both witnesses come
out at the counts they should — 151 / 57 / 94 and 140 / 51 / 89 — and Lenten
Espagnole picks up the parent it states, because the self-reference that
suppressed it is gone either way. What moves is which identity a parent
names:&lt;/p&gt;
&lt;pre&gt;&lt;code class="language-diff"&gt;  escoffier-1909, binding on exact names only
- mornay-sauce  parent  bechamel
+ mornay-sauce  parent  bechamel-sauce
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Eight parent values move like that, four in each witness: in 1909, Cream
Sauce and Mornay off &lt;code&gt;bechamel&lt;/code&gt;, Maltese and Mousseline off &lt;code&gt;hollandaise&lt;/code&gt;;
in 1907, Cream Sauce off &lt;code&gt;bechamel&lt;/code&gt;, and Maltese, Mousseline and Noisette
off &lt;code&gt;hollandaise&lt;/code&gt;. Mornay is the asymmetry — the 1907 scan splits its
heading into &lt;code&gt;MORN AY SAUCE&lt;/code&gt; and its first input reads &lt;code&gt;Bdchamel Sauce&lt;/code&gt;,
which reaches no catalogued name, so that witness has no &lt;code&gt;bechamel&lt;/code&gt; parent
there to move, as
&lt;a href="https://alberto.codes/blog/2026-09-05-the-parent-finally-has-a-verb"&gt;the parent post&lt;/a&gt; worked
through. &lt;code&gt;saucier tree hollandaise&lt;/code&gt; then prints a bare heading with no
children in either witness, and the family survives only under
&lt;code&gt;hollandaise-sauce&lt;/code&gt;, which is a heading and not a mother. That is ADR-0008's
coalescing rule — names reaching one preparation coalesce under the mother
concept — quietly reversed, and it is the rule
&lt;a href="https://alberto.codes/blog/2026-08-21-marrow-sauce-finally-has-a-parent"&gt;the Marrow Sauce post&lt;/a&gt;
spent two and a half thousand words getting right. Eight records that were
reading the book correctly would have paid for one that was not.&lt;/p&gt;
&lt;p&gt;The rule the proposal was aimed at turns out to decide almost nothing. Ranking by
source order arbitrates two of the ten bindings, and both of them are
velouté, the one mother in this book with three names carrying its word:&lt;/p&gt;
&lt;pre&gt;&lt;code class="language-text"&gt;veloute, escoffier-1909   ORDINARY VELOUTÉ SAUCE
                          VELOUTÉ DE VOLAILLE
                          ALLEMANDE SAUCE OR THICKENED VELOUTÉ
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Seven of the other eight bindings have a single candidate each, and the
1909 Espagnole binding never reaches the candidate list because its exact
name wins first. Espagnole in 1907 had a single candidate too. Nothing
was ranked, so no ranking rule could have prevented it. Measuring the
proposal is what made that visible, and the measurement is the reason the
actual defect got named: the run-match branch
never tested whether the name it crowned belonged to the base or to a
derivative of the base.&lt;/p&gt;
&lt;h2&gt;The sentence a chef would nod at&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;A mother does not bind by a name run to a preparation whose opening
paragraph states that mother.&lt;/strong&gt; An entry that says it is made from Espagnole
is not Espagnole.&lt;/p&gt;
&lt;p&gt;That is
&lt;a href="https://github.com/Alberto-Codes/saucier/blob/1664e75e95c6806ce604cbab0436e4f1bdaab1e0/docs/adr/0018-a-mother-does-not-bind-to-its-derivative.md"&gt;ADR-0018&lt;/a&gt;,
and the test it applies is not new. It is the statement test ADR-0008
already used to read parents out of prose: a whole run of words, inside one
sentence, of the opening paragraph. Two questions now share it, so it moved
into
&lt;a href="https://github.com/Alberto-Codes/saucier/blob/1664e75e95c6806ce604cbab0436e4f1bdaab1e0/src/saucier/domain/statement.py"&gt;&lt;code&gt;domain/statement.py&lt;/code&gt;&lt;/a&gt;
where neither caller owns it. Resolution asks which preparations an opening
states. The lookup asks whether a candidate states the base it is being
offered as. Same reading, opposite direction.&lt;/p&gt;
&lt;p&gt;The opening paragraph is the right place to look because of how Escoffier
writes. The first paragraph of an entry is its ingredient list. A sauce
named there is a sauce this one is built from, not one it is being compared
against — entry 22's own third paragraph, &lt;code&gt;The time required for the despumation of an Espagnole&lt;/code&gt; at line 1755, is not a claim of derivation, and
the guard does not read that far.&lt;/p&gt;
&lt;p&gt;An exact catalogued name still wins outright, before the guard runs. A base
that names itself exactly in its own heading keeps its own identity; the
guard only ever removes run matches. Ordering is untouched, and the survivors
still rank by source order. When nothing survives, the mother is
uncatalogued in that witness, which is the honest outcome and not a
fallback:&lt;/p&gt;
&lt;pre&gt;&lt;code class="language-console"&gt;$ uv run saucier show espagnole --source escoffier-1907
no preparation named 'espagnole'
[exit 1]
&lt;/code&gt;&lt;/pre&gt;
&lt;h2&gt;One line of JSON&lt;/h2&gt;
&lt;pre&gt;&lt;code class="language-console"&gt;$ uv run saucier tree espagnole --source escoffier-1907
espagnole  [espagnole]
├── HALF GLAZE  (en)
│   ├── SAUCE BORDELAISE  (fr)
│   ├── BROWN CHAUD=FROID SAUCE  (en)
│   ├── DEVILLED SAUCE  (en)
│   ├── ITALIAN SAUCE  (en)
│   ├── LYONNAISE SAUCE  (en)
│   ├── MADEIRA SAUCE  (en)
│   ├── PERIQUEUX SAUCE  (en)
│   ├── PIQUANTE SAUCE  (en)
│   └── ROBERT SAUCE  (en)
├── LENTEN ESPAGNOLE  (fr)
├── ORDINARY POIVRADE SAUCE  (en)
└── POIVRADE SAUCE FOR VENISON  (en)
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;The tree heads on the bare concept, because the witness carries the
derivations without carrying a usable name for what they derive from, and
the twelve preparations that were beneath Lenten Espagnole are beside it
now. &lt;code&gt;tree lenten-espagnole&lt;/code&gt; prints one line, &lt;code&gt;derives from espagnole&lt;/code&gt;, and
no children.&lt;/p&gt;
&lt;p&gt;The whole change to the parsed data is one line:&lt;/p&gt;
&lt;pre&gt;&lt;code class="language-diff"&gt;  data/escoffier-1907.json, line 126
-       &amp;quot;parent&amp;quot;: null,
+       &amp;quot;parent&amp;quot;: &amp;quot;espagnole&amp;quot;,
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;That is entry 24 recording the base its first sentence names. It is the
only parent value that moves in either witness, and it is the 51st derived
sauce in the 1907 catalogue:&lt;/p&gt;
&lt;pre&gt;&lt;code class="language-console"&gt;escoffier-1909  151 sauces, 57 derived, 94 unresolved
escoffier-1907  140 sauces, 51 derived, 89 unresolved
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;140 / 50 / 90 became 140 / 51 / 89. The 1909 census does not move, and its
JSON and JSONL come out byte-identical to the previous release — the guard
does remove one 1909 candidate, &lt;code&gt;VELOUTÉ DE VOLAILLE&lt;/code&gt;, whose opening states
velouté, but it was already losing to entry 25 on source order, so nothing
downstream notices. Two candidates removed across both witnesses, one
binding changed, nine correct bindings untouched. The &lt;code&gt;diff&lt;/code&gt; summary moves
from &lt;code&gt;11 unmatched, 19 parent-changed, 36 ocr-suspected&lt;/code&gt; to &lt;code&gt;11 unmatched, 18 parent-changed, 35 ocr-suspected&lt;/code&gt;. One row leaves and both counters drop,
because that row carried both labels: it is Lenten Espagnole's own, marked
&lt;code&gt;parent-changed, ocr-suspected&lt;/code&gt; and reading
&lt;code&gt;lenten-espagnole (none) / espagnole&lt;/code&gt;. The 1909 witness always resolved that
parent, and the two printings now agree on it.&lt;/p&gt;
&lt;p&gt;ADR-0008 stays accepted with its mother-binding clause amended, and
everything else in it intact. Its subject, shadow, ambiguity, and cycle
rules behave as before. Names reaching one preparation still coalesce under
the mother concept, so Mornay still reads &lt;code&gt;parent bechamel&lt;/code&gt; and
&lt;a href="https://alberto.codes/blog/2026-08-21-marrow-sauce-finally-has-a-parent"&gt;the record of that edge&lt;/a&gt;
still stands. A fix that is one sentence long and moves one field is what
you get when the sentence is about the book rather than about the code.&lt;/p&gt;
&lt;h2&gt;Where this guard runs out&lt;/h2&gt;
&lt;p&gt;It reads the witness's opening prose, so it is only as good as that prose
survived. In 1907 the &lt;code&gt;VELOUTE DE VOLAILLE&lt;/code&gt; opening does not state velouté,
although the 1909 opening does — which is why the guard removes that
candidate in one witness and not the other. Damage to a base's heading &lt;em&gt;and&lt;/em&gt;
to a derivative's opening paragraph would defeat it exactly the way the
heading alone defeated the old rule.&lt;/p&gt;
&lt;p&gt;The guard also reads the true base's own opening, and a base that stated its
own name there would remove itself and hand the heading to the next match in
source order. None do, which is what it means that the eight name-run
bindings survived the test: Escoffier opens a base with its quantities, not
with its name. The 1909 &lt;code&gt;ESPAGNOLE&lt;/code&gt; binding is an exact catalogued name and
returns before the guard runs, so it never faced the test at all.&lt;/p&gt;
&lt;p&gt;Entry 22 comes within a paragraph of a different trap. It names Espagnole at
line 1755, in its third paragraph, discussing how long despumation takes, and
nothing would hold it back from stating itself: an uncatalogued mother is
keyed by a bare concept that no entry holds, so no entry is held out from
stating it. Had that paragraph opened the entry, though, entry 22 would have
recorded no parent at all. Three lines on, the same paragraph names a second
mother — &lt;code&gt;the Mirepoix and the tomato are inserted from the first&lt;/code&gt;, at line
1758 — and ADR-0008 takes exactly one stated candidate or no parent, so two
of them resolve to none. What catches the near-miss is a rule that already
refuses to guess, not the cycle check: nothing reaches the cycle check, and
it reads heading lines, so it could not have caught a parent that names no
entry.&lt;/p&gt;
&lt;p&gt;And the crown is the only thing that got fixed. The scan still reads
&lt;code&gt;ESPAQNOLE&lt;/code&gt;, entry 22 is still catalogued as &lt;code&gt;brown-sauce&lt;/code&gt;, and
&lt;code&gt;tree brown-sauce --source escoffier-1907&lt;/code&gt; still prints one line with
nothing under it, while thirteen preparations hang off a concept whose
heading the witness cannot supply.
&lt;a href="https://github.com/Alberto-Codes/saucier/blob/v0.3.0/docs/adr/0013-repair-structure-never-content.md"&gt;ADR-0013&lt;/a&gt;
repairs the punctuation that delimits a record, never the characters inside
one, and one Q is a character inside one.
&lt;a href="https://github.com/Alberto-Codes/saucier/blob/v0.3.0/docs/adr/0014-a-damaged-witness-cannot-establish-absence.md"&gt;ADR-0014&lt;/a&gt;
says a damaged witness cannot establish absence. This says the neighbouring
thing: a damaged witness cannot establish identity by promoting a stated
derivative when the base's name disappears. Neither one repairs anything.
They both decline to conclude.&lt;/p&gt;
&lt;p&gt;Everything here reproduces from the tag:&lt;/p&gt;
&lt;pre&gt;&lt;code class="language-console"&gt;$ git clone https://github.com/Alberto-Codes/saucier
$ cd saucier &amp;amp;&amp;amp; git checkout v0.7.0
$ uv sync &amp;amp;&amp;amp; uv run saucier parse
$ uv run saucier diff escoffier-1907 escoffier-1909
$ uv run saucier tree espagnole --source escoffier-1907
$ uv run saucier show espaqnole --source escoffier-1907
$ uv run saucier show lenten-espagnole --source escoffier-1907
$ sed -n 1730p corpus/escoffier-1907.txt
$ sed -n 1795,1800p corpus/escoffier-1907.txt
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;If a heading in either witness is being read as something it is not, or if
the guard rejects a candidate you think belongs,
&lt;a href="https://github.com/Alberto-Codes/saucier/issues/new?template=extraction.yml"&gt;the issue template&lt;/a&gt;
asks for the entry number and the source lines. Line 1730 is where I would
start.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;The release:&lt;/strong&gt; &lt;a href="https://github.com/Alberto-Codes/saucier/releases/tag/v0.7.0"&gt;saucier v0.7.0&lt;/a&gt;
— the guard: a mother no longer binds by a name run to a preparation whose
opening paragraph states that mother, with ADR-0008's statement test moved
into the shared &lt;code&gt;saucier.domain.statement&lt;/code&gt; so both callers apply the same
one. In the 1907 scan the mother stays uncatalogued rather than falling back
to a derivative, &lt;code&gt;saucier show espagnole --source escoffier-1907&lt;/code&gt; refuses
and exits 1 where it used to print a Lenten sauce, and one field of the
parsed data moves. MIT.&lt;/p&gt;
</content:encoded>
    </item>
    <item>
      <title>Serve Gemma 4 31B on a 24 GiB card with the context it was packed for</title>
      <link>https://alberto.codes/blog/2026-09-06-serve-gemma-4-31b-on-a-24-gib-card-with-the-context-it-was-packed-for</link>
      <guid isPermaLink="true">https://alberto.codes/blog/2026-09-06-serve-gemma-4-31b-on-a-24-gib-card-with-the-context-it-was-packed-for</guid>
      <pubDate>Sun, 06 Sep 2026 00:00:00 GMT</pubDate>
      <description>Download the published fit24gib pack, install llama.cpp at the build the numbers were measured on, and serve it at 86,016 tokens of text context or 73,728 with an image aboard. Every flag comes from the model card or the vramfit repo's own serve how-to, and the section at the end says what to do when your card does not reach the boundary.</description>
      <content:encoded>&lt;h2&gt;Who this is for&lt;/h2&gt;
&lt;p&gt;You have a 24 GiB card and you want to run Gemma 4 31B on it with as
much context as the card will hold. You have read, or do not care
about, the argument for why a 14.92 GiB pack beats the 16.44 GiB
official build on this card. You want the server up.&lt;/p&gt;
&lt;p&gt;Assumed: a Linux box with a 24 GiB NVIDIA card, a working driver, and
a terminal. Not assumed: any quantization background. The why is
&lt;a href="https://alberto.codes/blog/2026-09-02-googles-4-bit-gemma-already-fit-my-card"&gt;the explanation post&lt;/a&gt;,
and the evidence is on
&lt;a href="https://huggingface.co/Alberto-Codes/gemma-4-31B-it-fit24gib-GGUF"&gt;the model card&lt;/a&gt;.
This post repeats none of it.&lt;/p&gt;
&lt;p&gt;Every number below is published, and the serve boundaries were
measured on an RTX 4090 running llama.cpp b10362 on 2026-08-31.
Numbers from another frame, the H100 build, the multi-image ladder,
say so where they appear. Your card is a different box, and the last
section is about what to do when a number does not reproduce.&lt;/p&gt;
&lt;h2&gt;Before you start&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Disk.&lt;/strong&gt; The decoder is 14.92 GiB and the projector sidecar is
629 MiB. The llama.cpp build or tarball adds a few GiB. Call it
20 GiB free.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Two files, one artifact.&lt;/strong&gt; The decoder alone serves text. Images
need the sidecar too. Both download in one command below.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;The license.&lt;/strong&gt; The weights carry the
&lt;a href="https://ai.google.dev/gemma/docs/gemma_4_license"&gt;Gemma 4 license note&lt;/a&gt;.
Read it before you serve them to anyone but yourself.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;The build is part of the claim.&lt;/strong&gt; The boundaries on the card were
measured at llama.cpp b10362. A newer build allocates the KV cache
its own way and may land a rung higher or lower. Pin the build first,
reproduce the boundary, then move if you want to.&lt;/p&gt;
&lt;h2&gt;1. Download the pack&lt;/h2&gt;
&lt;p&gt;The repo is public. No token needed.&lt;/p&gt;
&lt;pre&gt;&lt;code class="language-bash"&gt;uv tool install &amp;quot;huggingface_hub[cli]&amp;quot;
mkdir -p ~/models &amp;amp;&amp;amp; cd ~/models
hf download Alberto-Codes/gemma-4-31B-it-fit24gib-GGUF \
  gemma-4-31B-it-fit24gib.gguf gemma-4-31B-it-mmproj-q4km.gguf \
  --local-dir .
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;What lands:&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;File&lt;/th&gt;
&lt;th&gt;Bytes&lt;/th&gt;
&lt;th&gt;Size&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;gemma-4-31B-it-fit24gib.gguf&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;16,015,862,144&lt;/td&gt;
&lt;td&gt;14.92 GiB&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;gemma-4-31B-it-mmproj-q4km.gguf&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;659,537,504&lt;/td&gt;
&lt;td&gt;629 MiB&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;Check the decoder before you trust it. The card publishes the hash:&lt;/p&gt;
&lt;pre&gt;&lt;code class="language-bash"&gt;sha256sum gemma-4-31B-it-fit24gib.gguf
# 2a7bd7a7be6979c858258618ab576db573a7b671b45ee5e9785247341b8c3b1e
&lt;/code&gt;&lt;/pre&gt;
&lt;h2&gt;2. Install llama.cpp b10362&lt;/h2&gt;
&lt;p&gt;Two ways. The first is the same build number and the same backend the
card's boundaries were measured on. The second is the path the repo's
rented-H100 how-to takes.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Prebuilt Vulkan tarball.&lt;/strong&gt; The
&lt;a href="https://github.com/ggml-org/llama.cpp/releases/tag/b10362"&gt;b10362 release&lt;/a&gt;
ships a Linux Vulkan build:&lt;/p&gt;
&lt;pre&gt;&lt;code class="language-bash"&gt;cd ~
curl -LO https://github.com/ggml-org/llama.cpp/releases/download/b10362/llama-b10362-bin-ubuntu-vulkan-x64.tar.gz
tar xzf llama-b10362-bin-ubuntu-vulkan-x64.tar.gz
export BIN=~/llama-b10362    # the tarball unpacks flat into this directory
&amp;quot;$BIN/llama-server&amp;quot; --version
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;&lt;strong&gt;Build from source with CUDA.&lt;/strong&gt; The release ships no Linux CUDA
tarball, so CUDA means a build. Pin the commit the tag points at:&lt;/p&gt;
&lt;pre&gt;&lt;code class="language-bash"&gt;git clone https://github.com/ggml-org/llama.cpp ~/llama.cpp
cd ~/llama.cpp
git fetch --tags origin
git checkout 4801e3c56    # the commit tag b10362 points at
cmake -B build -DGGML_CUDA=ON -DCMAKE_BUILD_TYPE=Release
cmake --build build -j
export BIN=~/llama.cpp/build/bin
&amp;quot;$BIN/llama-server&amp;quot; --version
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Either way the version line reads &lt;code&gt;version: 10362 (4801e3c56)&lt;/code&gt;. On
the rented H100 that build took 429 seconds. CUDA is a different
runtime from Vulkan, so treat the boundaries below as the rung to
test first, not a promise. Step 5 shows how.&lt;/p&gt;
&lt;h2&gt;3. Serve text at the measured boundary&lt;/h2&gt;
&lt;p&gt;Note what is free before you load. The card's ladders ran with
23,629 to 23,631 MiB free on a 24,564 MiB device, under a desktop.&lt;/p&gt;
&lt;pre&gt;&lt;code class="language-bash"&gt;nvidia-smi --query-gpu=memory.total,memory.free --format=csv
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Then the command from the card, verbatim:&lt;/p&gt;
&lt;pre&gt;&lt;code class="language-bash"&gt;M=~/models
&amp;quot;$BIN/llama-server&amp;quot; -m &amp;quot;$M/gemma-4-31B-it-fit24gib.gguf&amp;quot; \
  -c 86016 -ngl 99 -np 1 --port 8991 &amp;gt; server.log 2&amp;gt;&amp;amp;1 &amp;amp;
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Three flags carry the claim.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;code&gt;-c 86016&lt;/code&gt; is the measured text boundary. The next rung, 90,112,
fails to load.&lt;/li&gt;
&lt;li&gt;&lt;code&gt;-ngl 99&lt;/code&gt; offloads every layer. Any layer left on the CPU frees
VRAM and invalidates the comparison.&lt;/li&gt;
&lt;li&gt;&lt;code&gt;-np 1&lt;/code&gt; is one slot. The b10362 server defaults to four, which on
this geometry adds about 2,400 MiB of sliding-window cache and
fails loads that fit at one.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Wait for &lt;code&gt;model loaded&lt;/code&gt; in the log, then hit the health route. This
build logs &lt;code&gt;all slots are idle&lt;/code&gt; only at trace verbosity, so do not
wait on it: a check on that line waits forever with the server
healthy.&lt;/p&gt;
&lt;pre&gt;&lt;code class="language-bash"&gt;while pgrep -x llama-server &amp;gt; /dev/null &amp;amp;&amp;amp; ! grep -q &amp;quot;model loaded&amp;quot; server.log; do sleep 2; done
curl -s localhost:8991/health
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;The reply is &lt;code&gt;{&amp;quot;status&amp;quot;:&amp;quot;ok&amp;quot;}&lt;/code&gt;. If the loop returns before
&lt;code&gt;model loaded&lt;/code&gt; appears, the server exited, and the tail of
&lt;code&gt;server.log&lt;/code&gt; says why; see &amp;quot;If it does not fit&amp;quot; below.&lt;/p&gt;
&lt;p&gt;A loaded server is not a server on the GPU. Confirm the offload
before you trust any number below:&lt;/p&gt;
&lt;pre&gt;&lt;code class="language-bash"&gt;grep -i offloaded server.log
grep -i &amp;quot;gpu-layers option will be ignored&amp;quot; server.log
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;The first prints &lt;code&gt;offloaded N/N layers to GPU&lt;/code&gt;, both numbers equal
and neither of them zero. The second prints nothing. &lt;code&gt;offloaded 0/N&lt;/code&gt;
means &lt;code&gt;-ngl 99&lt;/code&gt; was parsed and then discarded, and the second grep
says why: the build has no backend for your card, a CPU-only archive
or a driver it cannot see. Nothing else looks wrong — the server
loads, the health route answers &lt;code&gt;ok&lt;/code&gt;, requests come back — and every
layer runs on the CPU at a fraction of the speed, while every
boundary in this post assumes the card is doing the work. Fix step 2
before you measure anything.&lt;/p&gt;
&lt;h2&gt;4. Send one request&lt;/h2&gt;
&lt;p&gt;The server speaks the OpenAI chat shape. One text turn:&lt;/p&gt;
&lt;pre&gt;&lt;code class="language-bash"&gt;curl -s localhost:8991/v1/chat/completions \
  -H 'Content-Type: application/json' \
  -d '{
    &amp;quot;messages&amp;quot;: [{&amp;quot;role&amp;quot;: &amp;quot;user&amp;quot;, &amp;quot;content&amp;quot;: &amp;quot;In one sentence, what is a sliding-window attention layer?&amp;quot;}],
    &amp;quot;max_tokens&amp;quot;: 64
  }' | python3 -c 'import json,sys; print(json.load(sys.stdin)[&amp;quot;choices&amp;quot;][0][&amp;quot;message&amp;quot;][&amp;quot;content&amp;quot;])'
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;That answers from inside the 86,016-token envelope. It does not tell
you what a full envelope costs to decode: the card's boundary check
decoded five tokens, and throughput at the boundary is unmeasured.
The throughput the card does publish was taken at 8,192 tokens of
context on the same 4090, 47.8 tokens per second for this pack
against 43.3 for Google's Q4_0.&lt;/p&gt;
&lt;h2&gt;5. Read the VRAM numbers back&lt;/h2&gt;
&lt;p&gt;While the server is up:&lt;/p&gt;
&lt;pre&gt;&lt;code class="language-bash"&gt;nvidia-smi --query-gpu=memory.used,memory.free --format=csv
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;At the 86,016 boundary on the 4090 the load passed with 143 MiB
free. That is what a fit bar looks like: the card is full, and the
next rung is the one that fails. If you see a few gigabytes free, one
of the three flags above is not doing what you think, most often a
default &lt;code&gt;-np 4&lt;/code&gt; from a wrapper script.&lt;/p&gt;
&lt;p&gt;To confirm the boundary is real on your box, stop the server and load
one rung higher:&lt;/p&gt;
&lt;pre&gt;&lt;code class="language-bash"&gt;pkill -x llama-server
until ! pgrep -x llama-server &amp;gt; /dev/null; do sleep 1; done
&amp;quot;$BIN/llama-server&amp;quot; -m &amp;quot;$M/gemma-4-31B-it-fit24gib.gguf&amp;quot; -c 90112 -ngl 99 -np 1 --port 8991
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;If it fails, you have confirmed the boundary; if it loads, stop it
with Ctrl-C and keep climbing a rung at a time until one fails; that
rung is yours. A higher rung means your box idles with more VRAM free
than the card's did. The card
calls the tuple of box, build, backend, and free VRAM before load the
frame, and it prints the frame beside every boundary, because the
same file served 81,920 for text three days earlier in a frame with
less idle VRAM. Write yours down the same way. It is the only way two
boundaries compare.&lt;/p&gt;
&lt;h2&gt;6. Serve images&lt;/h2&gt;
&lt;p&gt;Images need the sidecar and two more flags. Stop whichever server is
still holding the card, start the image server from the card, and wait
for &lt;code&gt;model loaded&lt;/code&gt; the same way as in step 3:&lt;/p&gt;
&lt;pre&gt;&lt;code class="language-bash"&gt;pkill -x llama-server
until ! pgrep -x llama-server &amp;gt; /dev/null; do sleep 1; done
&amp;quot;$BIN/llama-server&amp;quot; -m &amp;quot;$M/gemma-4-31B-it-fit24gib.gguf&amp;quot; \
  --mmproj &amp;quot;$M/gemma-4-31B-it-mmproj-q4km.gguf&amp;quot; \
  -c 73728 -ngl 99 -np 1 --mtmd-batch-max-tokens 264 \
  --port 8991 &amp;gt; server.log 2&amp;gt;&amp;amp;1 &amp;amp;
while pgrep -x llama-server &amp;gt; /dev/null &amp;amp;&amp;amp; ! grep -q &amp;quot;model loaded&amp;quot; server.log; do sleep 2; done
&lt;/code&gt;&lt;/pre&gt;
&lt;ul&gt;
&lt;li&gt;&lt;code&gt;-c 73728&lt;/code&gt; is the measured one-image boundary. The next rung,
77,824, loads and then fails at encode time.&lt;/li&gt;
&lt;li&gt;&lt;code&gt;--mtmd-batch-max-tokens 264&lt;/code&gt; caps the encode batch at one
1280×720 image, which is 264 image tokens. Without it the server
packs up to 1,024 image tokens into one encode graph, two images
share a graph, and the graph asks for 328 MiB against a 150.63 MiB
one-image reserve. On 2026-09-02 that crashed the server on the
second image at both configurations. With the cap, the same ladder
filled the window to a clean context refusal.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Keep about 200 MiB free beyond the load. The image encode allocates at
request time, and this build crashes on that failure instead of
refusing.&lt;/p&gt;
&lt;p&gt;An image request is the same chat shape with an &lt;code&gt;image_url&lt;/code&gt; part
carrying a base64 data URL. Write it to a file first: a base64
screenshot is bigger than the single-argument limit a Linux shell
allows, and an inline &lt;code&gt;-d&lt;/code&gt; fails with &amp;quot;Argument list too long&amp;quot;.&lt;/p&gt;
&lt;pre&gt;&lt;code class="language-bash"&gt;python3 - &amp;lt;&amp;lt;'PY'
import base64, json
img = base64.b64encode(open(&amp;quot;screenshot.png&amp;quot;, &amp;quot;rb&amp;quot;).read()).decode()
json.dump({&amp;quot;max_tokens&amp;quot;: 64, &amp;quot;messages&amp;quot;: [{&amp;quot;role&amp;quot;: &amp;quot;user&amp;quot;, &amp;quot;content&amp;quot;: [
    {&amp;quot;type&amp;quot;: &amp;quot;text&amp;quot;, &amp;quot;text&amp;quot;: &amp;quot;What is on this screen?&amp;quot;},
    {&amp;quot;type&amp;quot;: &amp;quot;image_url&amp;quot;, &amp;quot;image_url&amp;quot;: {&amp;quot;url&amp;quot;: &amp;quot;data:image/png;base64,&amp;quot; + img}},
]}]}, open(&amp;quot;req.json&amp;quot;, &amp;quot;w&amp;quot;))
PY
curl -s localhost:8991/v1/chat/completions -H 'Content-Type: application/json' \
  -d @req.json | python3 -c 'import json,sys; print(json.load(sys.stdin)[&amp;quot;choices&amp;quot;][0][&amp;quot;message&amp;quot;][&amp;quot;content&amp;quot;])'
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;One 768×768 image costs 256 decoder tokens on this pack, measured at
the server. The 1280×720 screenshot above is 264 image tokens, 271
once its wrapper is counted, which is why the cap is 264 and the
decoder sees 271.&lt;/p&gt;
&lt;h2&gt;What you should see&lt;/h2&gt;
&lt;p&gt;The card's serve ladders, measured 2026-08-31, RTX 4090, llama.cpp
b10362 Vulkan, &lt;code&gt;-ngl 99 -np 1&lt;/code&gt;, KV cache f16, 4,096-token rungs:&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Serving shape&lt;/th&gt;
&lt;th&gt;This pack&lt;/th&gt;
&lt;th&gt;Google's QAT Q4_0&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Text only, max load&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;86,016&lt;/strong&gt; (fails at 90,112)&lt;/td&gt;
&lt;td&gt;65,536 (fails at 69,632)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;One image aboard, max load&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;73,728&lt;/strong&gt; (encode fails at 77,824)&lt;/td&gt;
&lt;td&gt;49,152 (encode fails at 53,248)&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;The gain is 20,480 tokens of text and 24,576 with an image, on that
frame. Two of the three things that set a boundary, the build and the
box, are yours.&lt;/p&gt;
&lt;p&gt;Quality is not the subject of this post. The card's five held-out
benchmarks put this pack at four ties and one win against Google's
build. Read the tables there, not here.&lt;/p&gt;
&lt;h2&gt;If it does not fit&lt;/h2&gt;
&lt;p&gt;Work down this list in order. Each item is a boundary the card
already crossed.&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;&lt;strong&gt;Drop one rung.&lt;/strong&gt; The rungs are 4,096 tokens. Try 81,920 for
text, 69,632 with an image. The 81,920 text rung reproduced on
2026-08-31 with 465 MiB free at load, so it is the safer first
stop under a heavier desktop.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Check &lt;code&gt;-np&lt;/code&gt;.&lt;/strong&gt; A wrapper that sets slots for you costs about
2,400 MiB of sliding-window cache on this geometry. One slot, or
nothing here holds.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Check what else holds the card.&lt;/strong&gt; The ladders ran with
23,629 to 23,631 MiB free before each load. A rung is about
320 MiB on this geometry, so a browser or a second model can cost
one.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Serving images: keep the encode-batch cap.&lt;/strong&gt; Every crash on the
card's multi-image ladder traced to two images sharing one encode
graph. &lt;code&gt;--mtmd-batch-max-tokens 264&lt;/code&gt; is the fix, not a tuning
knob.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;A different backend is a different frame.&lt;/strong&gt; CUDA, ROCm, or a
newer llama.cpp allocates differently. Measure your own ladder
with step 5 and report the frame with the number.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;If you have more card than this, the repo has
&lt;a href="https://github.com/Alberto-Codes/vramfit/blob/main/docs/how-to/serve-gemma-4-fit24gib-on-a-rented-h100.md"&gt;a how-to for serving the pack on a rented H100&lt;/a&gt;.
It is the instrument behind the card's real-GUI campaign: a CUDA
build of the same b10362 tag, the same two files, and context 8,192
for single-image evaluation. The 24 GiB boundaries and the
encode-batch flag belong to the 4090 frame and do not carry over.&lt;/p&gt;
&lt;p&gt;What you have at the end is not the card's number. It is a boundary
measured on your own box, with its frame written beside it, which is
the only kind of number the card ever claimed.&lt;/p&gt;
&lt;h2&gt;Where the numbers come from&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;The pack and its card:&lt;/strong&gt;
&lt;a href="https://huggingface.co/Alberto-Codes/gemma-4-31B-it-fit24gib-GGUF"&gt;gemma-4-31B-it-fit24gib-GGUF&lt;/a&gt;.
The serve commands, the ladders, the reproduction traps, the hashes.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;The map behind it:&lt;/strong&gt;
&lt;a href="https://huggingface.co/datasets/Alberto-Codes/gemma-4-31B-it-sensitivity-maps"&gt;gemma-4-31B-it-sensitivity-maps&lt;/a&gt;,
published 2026-09-04, so &lt;code&gt;vramfit plan&lt;/code&gt; can re-solve this model for
a different budget.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;The baseline:&lt;/strong&gt;
&lt;a href="https://huggingface.co/google/gemma-4-31B-it-qat-q4_0-gguf"&gt;google/gemma-4-31B-it-qat-q4_0-gguf&lt;/a&gt;,
16.44 GiB.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;The tool:&lt;/strong&gt;
&lt;a href="https://github.com/Alberto-Codes/vramfit/releases/tag/v0.5.0"&gt;vramfit v0.5.0&lt;/a&gt;
— the tag this guide serves from: the thin &lt;code&gt;gguf&lt;/code&gt; extra that reads a
pack without torch, &lt;code&gt;pack&lt;/code&gt; comparing its packed bytes against the
recipe's prediction, and every refusal raised under one root. MIT.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;The argument:&lt;/strong&gt;
&lt;a href="https://alberto.codes/blog/2026-09-02-googles-4-bit-gemma-already-fit-my-card"&gt;the explanation post&lt;/a&gt;.
The pipeline that built the file is
&lt;a href="https://alberto.codes/blog/2026-08-15-fit-a-model-to-the-gpu-you-actually-have"&gt;fit a model to the GPU you actually have&lt;/a&gt;.&lt;/li&gt;
&lt;/ul&gt;
</content:encoded>
    </item>
    <item>
      <title>The parent finally has a verb. The book still will not say how many minutes.</title>
      <link>https://alberto.codes/blog/2026-09-05-the-parent-finally-has-a-verb</link>
      <guid isPermaLink="true">https://alberto.codes/blog/2026-09-05-the-parent-finally-has-a-verb</guid>
      <pubDate>Sat, 05 Sep 2026 00:00:00 GMT</pubDate>
      <description>At v0.5.0 the saucier record for Mornay said it derives from Béchamel and stopped there, two lines above the sentence that says how. v0.6.0 records that sentence for one preparation, by hand, as six operations in the book's own words, with every number the text gives and every one it withholds left empty. The check that refuses a misquote also turned up the first confirmed difference between the 1907 and 1909 printings.</description>
      <content:encoded>&lt;p&gt;Here is what &lt;code&gt;saucier show mornay&lt;/code&gt; printed at &lt;code&gt;v0.5.0&lt;/code&gt;, cut to the two lines
that matter:&lt;/p&gt;
&lt;pre&gt;&lt;code class="language-console"&gt;  parent  bechamel

Boil one pint of Béchamel Sauce with one-quarter pint of the _fumet_
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;The record says Mornay derives from Béchamel. The sentence underneath, at
line 2439 of the 1909 text, two lines below the heading of entry 91, says
what is done with the Béchamel, and the record does not read it. &lt;a href="https://alberto.codes/blog/2026-08-21-marrow-sauce-finally-has-a-parent"&gt;The Marrow Sauce post&lt;/a&gt;
spent two and a half thousand words getting the parent edge right, and the edge it got
right is an input with the operation stripped off. A parent says what a sauce
is built from. It does not say boil, reduce, or finish, and it does not say
how much or how far.&lt;/p&gt;
&lt;p&gt;&lt;code&gt;v0.6.0&lt;/code&gt; reads that sentence, for one preparation, by hand. The census does
not move: 151 sauces, 57 derived, 94 unresolved in the 1909 text, and 140,
50, 90 in the 1907 scan, the same census as
&lt;a href="https://alberto.codes/blog/2026-09-04-i-cut-the-last-sauce-off-the-file"&gt;the stream post&lt;/a&gt;, whose
151 and 140 these split. What moves is the record underneath one of them.&lt;/p&gt;
&lt;h2&gt;Six operations, in the book's order&lt;/h2&gt;
&lt;pre&gt;&lt;code class="language-console"&gt;$ uv run saucier show mornay
MORNAY SAUCE
entry 91, line 2437, transcription of escoffier-1909
  term  MORNAY SAUCE  [en]  mornay-sauce
  parent  bechamel
  procedure  6 operations, recorded by hand
    Boil      Béchamel Sauce [fr] 1 pint, fumet [fr] 1/4 pint
    Reduce    criterion: by a good quarter (unresolved)
    add       Gruyère [fr] 2 oz., Parmesan [en] 2 oz.
    Put       duration: a few minutes (unresolved), on the fire again
    stirring  instrument: small whisk, criterion: the melting of the cheese (unresolved)
    Finish    butter [en] 2 oz., away from the fire, added by degrees
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;And the eight lines it was read from, lines 2439 to 2446 of the committed
corpus, exactly as Project Gutenberg transcribed them:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Boil one pint of Béchamel Sauce with one-quarter pint of the &lt;em&gt;fumet&lt;/em&gt;
of the fish, poultry, or vegetable, which is to constitute the dish.
Reduce by a good quarter, and add two oz. of Gruyère and two oz. of
grated Parmesan.&lt;/p&gt;
&lt;p&gt;Put the sauce on the fire again for a few minutes, and ensure the
melting of the cheese by stirring with a small whisk. Finish the sauce
away from the fire with two oz. of butter added by degrees.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;&lt;img src="https://alberto.codes/saucier-mornay-procedure.svg" alt="The Mornay entry and its procedure side by side. Left, lines 2439 to 2446 of escoffier-1909, the eight lines of the body, with six clauses highlighted in reading order: Boil one pint of Béchamel Sauce with one-quarter pint of the fumet, Reduce by a good quarter, add two oz. of Gruyère and two oz. of grated Parmesan, Put the sauce on the fire again for a few minutes, ensure the melting of the cheese by stirring with a small whisk, and Finish the sauce away from the fire with two oz. of butter added by degrees. Right, the six operations as saucier show prints them, one per row, each joined to its clause by an arrow. Slots that hold a number are drawn solid green: 1 pint, 1/4 pint, 2 oz. three times. Slots the text leaves without a number are drawn in dashed orange and labelled unresolved: by a good quarter, a few minutes, the melting of the cheese. Caption: every word on the right is a run of words on the left, and the command checks that before it prints." /&gt;&lt;/p&gt;
&lt;p&gt;Every line on the right is a run of words on the left. &lt;code&gt;Boil&lt;/code&gt; takes two
inputs, each with the quantity the text gives: one pint, one-quarter pint.
&lt;code&gt;Reduce&lt;/code&gt; takes nothing and carries a criterion, the words the reduction is
carried to. &lt;code&gt;add&lt;/code&gt; takes two cheeses at two ounces each. &lt;code&gt;Put&lt;/code&gt; carries a
duration and a constraint, &lt;code&gt;on the fire again&lt;/code&gt;. &lt;code&gt;stirring&lt;/code&gt; carries an
instrument and a criterion. &lt;code&gt;Finish&lt;/code&gt; takes the butter and two constraints,
&lt;code&gt;away from the fire&lt;/code&gt; and &lt;code&gt;added by degrees&lt;/code&gt;. That is the whole entry, and
nothing in it is inferred.
&lt;a href="https://github.com/Alberto-Codes/saucier/blob/v0.6.0/docs/adr/0017-a-procedure-quotes-its-witness.md"&gt;ADR-0017&lt;/a&gt;
is the record of the shape, and the glossary gains ten terms for it,
because this project uses one word per thing and this release needed ten
new things: procedure, operation, wording, input, parameter, criterion,
constraint, instrument, recorder, unrecorded.&lt;/p&gt;
&lt;h2&gt;Three slots the text left empty&lt;/h2&gt;
&lt;p&gt;Look at the three &lt;code&gt;(unresolved)&lt;/code&gt; marks. Béchamel gets &lt;code&gt;1 pint&lt;/code&gt;. Gruyère
gets &lt;code&gt;2 oz.&lt;/code&gt;. The reduction gets &lt;code&gt;by a good quarter&lt;/code&gt;, and no number. The
return to the fire gets &lt;code&gt;a few minutes&lt;/code&gt;, and no number. The stirring gets
&lt;code&gt;the melting of the cheese&lt;/code&gt;, which is a criterion with no number in it at
all.&lt;/p&gt;
&lt;p&gt;That is deliberate, and it is the same rule this series has applied to
parents since &lt;a href="https://alberto.codes/blog/2026-08-19-there-is-no-model-in-this-parser"&gt;the first post&lt;/a&gt;.
A parameter holds the words, the number the words give, and the unit they
name. &lt;code&gt;one-quarter pint&lt;/code&gt; records the fraction one over four and the unit
pint. &lt;code&gt;a few minutes&lt;/code&gt; records the unit minutes and no number, because the
text gives none. &lt;code&gt;by a good quarter&lt;/code&gt; names a degree and not a quantity, so
it records no number either. No code fills the slot. Anyone who has cooked a
Mornay knows roughly how many minutes &amp;quot;a few&amp;quot; is, and a model would happily
write 3. The parser writes nothing, because Escoffier wrote nothing, and
the moment this record holds a number the book does not the record stops
being the baseline a model has to beat.&lt;/p&gt;
&lt;p&gt;The stream post argued that &lt;code&gt;null&lt;/code&gt; in a &lt;code&gt;parent&lt;/code&gt; field means the source
declined to state one, never that the sauce has none. This is the first
time the same absence is recorded against something other than a parent. A
duration with no number is a fact about the text.&lt;/p&gt;
&lt;h2&gt;The record cannot say what the entry does not&lt;/h2&gt;
&lt;p&gt;The six operations are Python literals in an adapter, written by hand
against lines 2439 to 2446. A hand can misquote. So the command never
prints a procedure it has not first found in the body.&lt;/p&gt;
&lt;p&gt;Each operation carries its &lt;code&gt;wording&lt;/code&gt;, the whole clause it was read from.
Each input, criterion, duration, and constraint carries its own wording, and
the entity refuses an element whose words do not lie inside its operation's
wording.
Then, before &lt;code&gt;show&lt;/code&gt; prints anything, it collapses whitespace on both sides
and looks for each operation's wording in the body, in order, each one
after the last. An operation the body does not carry, or carries out of
order, is reported, and the command exits 2 with nothing on standard
output:&lt;/p&gt;
&lt;pre&gt;&lt;code class="language-console"&gt;$ uv run saucier show mornay   # with a hand record that says &amp;quot;by a good half&amp;quot;
saucier: MORNAY SAUCE at line 2437 of escoffier-1909 does not state 'Reduce by a good half'
[exit 2]
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;The first version of this printed the title, the reference, the terms, and
the parent before it ran the check, so a piped reader got half a record and
then a non-zero exit. Review caught it, and the check now runs before the
first byte. The stream post's first reader accepted a file missing its last
line. This release's first &lt;code&gt;show&lt;/code&gt; printed half a preparation beside a
refusal. Same lesson, one layer up: a check that runs after the output has
started is a comment.&lt;/p&gt;
&lt;p&gt;What the check proves is narrow, and the ADR says so. It proves the words
are there and in that order. It does not prove the reader parsed the clause
correctly. Three choices in the Mornay record are the reader's, and they
are written down so a disagreement has something to point at. &lt;code&gt;ensure the melting of the cheese by stirring with a small whisk&lt;/code&gt; records the verb as
&lt;code&gt;stirring&lt;/code&gt;, with the melting as the criterion and the whisk as the
instrument. &lt;code&gt;the sauce&lt;/code&gt; and &lt;code&gt;the cheese&lt;/code&gt; name the preparation in progress
and are not inputs. &lt;code&gt;Put the sauce on the fire again&lt;/code&gt; names no heat, so no
heat is recorded.&lt;/p&gt;
&lt;h2&gt;The scan keeps its verb and loses its parent&lt;/h2&gt;
&lt;p&gt;The 1907 witness carries the same entry at line 2864, and the release
records a second procedure for it, in the scan's words, because a procedure
quotes its witness and the two witnesses do not read alike:&lt;/p&gt;
&lt;pre&gt;&lt;code class="language-console"&gt;$ uv run saucier show morn-ay-sauce --source escoffier-1907
MORN AY SAUCE
entry 91, line 2864, ocr of escoffier-1907
  term  MORN AY SAUCE  [en]  morn-ay-sauce
  parent  (unresolved)
  stated  no candidate
  procedure  6 operations, recorded by hand
    Boil      Bdchamel Sauce [fr] 1 pint, fumet [fr] 1/4 pint
    Reduce    criterion: by a good quarter (unresolved)
    add       Gruy^re [fr] 2 oz., Parmesan [en] 2 oz.
    Put       duration: a few minutes (unresolved), on the fire again
    stirring  instrument: small whisk, criterion: the melting of the cheese (unresolved)
    Finish    butter [en] 2 oz., away from the fire, added by degrees
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;The heading reads &lt;code&gt;MORN AY SAUCE&lt;/code&gt;, so &lt;code&gt;show mornay&lt;/code&gt; finds nothing in the
scan. The first input reads &lt;code&gt;Bdchamel Sauce&lt;/code&gt;, whose folded form reaches no
catalogued name, so the scan's Mornay has no parent and no candidate. The
text resolves it. The verb is on the record in both. &lt;code&gt;Gruy^re&lt;/code&gt; stays
&lt;code&gt;Gruy^re&lt;/code&gt;, and the &lt;code&gt;Reduce&lt;/code&gt; clause carries the running page header &lt;code&gt;40 GUIDE TO MODERN COOKERY&lt;/code&gt; inside its wording, because the scan carries it there
and &lt;a href="https://github.com/Alberto-Codes/saucier/blob/v0.3.0/docs/adr/0013-repair-structure-never-content.md"&gt;ADR-0013&lt;/a&gt;
repairs the punctuation that delimits a record, never the characters inside
one. The Périgueux page break from the stream post is the same running header,
ten pages earlier, doing the same damage.&lt;/p&gt;
&lt;h2&gt;The one difference that is not the scanner&lt;/h2&gt;
&lt;p&gt;&lt;a href="https://alberto.codes/blog/2026-09-01-i-added-a-second-copy-of-the-same-book"&gt;The second-copy post&lt;/a&gt;
ended on a sentence I have repeated in every post since: not one confirmed
editorial difference between the 1907 and 1909 printings. Every candidate
was the scanner or the reader. The README said so at &lt;code&gt;v0.5.0&lt;/code&gt;.&lt;/p&gt;
&lt;p&gt;Writing the 1907 procedure by hand meant reading line 2867 of the scan
against line 2440 of the text, word by word, because the check would refuse
anything else. The fumet is &lt;code&gt;of that fish which is to constitute the dish&lt;/code&gt;
in 1907. It is &lt;code&gt;of the fish, poultry, or vegetable, which is to constitute the dish&lt;/code&gt; in 1909. No scanner adds two nouns and three commas. Between the first
printing and the revised edition the sentence widened, from fish to fish,
poultry, or vegetable. Whose hand widened it, the two texts do not say.&lt;/p&gt;
&lt;p&gt;That is one difference, confirmed by hand, on two lines anyone can open.
The diff command has still confirmed none, and the README now says exactly
that: the diff has confirmed nothing, and one difference has been confirmed
by a reader with both texts in front of them. It took recording a
procedure to find it, because a procedure quotes the witness and a parent
only names it. &lt;code&gt;bechamel&lt;/code&gt; is the same concept in both books. &lt;code&gt;of that fish&lt;/code&gt; and &lt;code&gt;of the fish, poultry, or vegetable&lt;/code&gt; are not the same words.&lt;/p&gt;
&lt;h2&gt;What I am not claiming&lt;/h2&gt;
&lt;p&gt;One preparation is recorded, once per witness, and a test pins the count
at one. Two procedures written by hand is not extraction. The rule that
would read a second one out of an entry does not exist yet, and until it
does the count stays where it is. What does exist is the port the hand
record enters through, and the line &lt;code&gt;recorded by hand&lt;/code&gt; in the output is
the recorder naming itself. A rule reader, or a model, is another
implementation behind the same port, and whatever it records will say who
read it. Cardinal Sauce, entry 69, still boils
Béchamel and finishes with lobster butter and still records no parent,
because two catalogued names sit in its opening paragraph and
&lt;a href="https://github.com/Alberto-Codes/saucier/blob/v0.3.0/docs/adr/0012-a-resolver-may-refuse-never-rank.md"&gt;a resolver may refuse, never rank&lt;/a&gt;.
The verb that would tell those two names apart, one boiled and one added at
the finish, is exactly what a procedure carries. Reading it by rule is the
record after this one, and the ten sauces
&lt;a href="https://alberto.codes/blog/2026-09-03-the-book-spells-it-out-at-line-1437"&gt;the line-1437 post&lt;/a&gt;
lost to a butter are waiting on it.&lt;/p&gt;
&lt;p&gt;The procedure is not stored and the interchange does not carry it. &lt;code&gt;show&lt;/code&gt;
fetches it beside the preparation, checks it, and prints it. The JSON files
under &lt;code&gt;data/&lt;/code&gt; did not change, &lt;code&gt;saucier/1&lt;/code&gt; did not change, and a consumer of
the stream sees no operation.
&lt;a href="https://github.com/Alberto-Codes/saucier/blob/v0.5.0/docs/adr/0016-jsonl-is-the-interchange-not-a-store.md"&gt;ADR-0016&lt;/a&gt;
says a new field earns a new schema version, and one preparation does not
earn one.&lt;/p&gt;
&lt;p&gt;And the &lt;code&gt;parent&lt;/code&gt; field is untouched. The procedure sits beside it and never
writes it. Mornay's first operation boils Béchamel, which is its parent, and
that agreement is a check on the resolver, not a replacement for it.&lt;/p&gt;
&lt;p&gt;Everything here reproduces from the tag:&lt;/p&gt;
&lt;pre&gt;&lt;code class="language-console"&gt;$ git clone https://github.com/Alberto-Codes/saucier
$ cd saucier &amp;amp;&amp;amp; git checkout v0.6.0
$ uv sync &amp;amp;&amp;amp; uv run saucier parse
$ uv run saucier show mornay
$ uv run saucier show morn-ay-sauce --source escoffier-1907
$ sed -n 2437,2446p corpus/escoffier-1909.txt
$ sed -n 2864,2878p corpus/escoffier-1907.txt
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Six operations, five quantities the book gives, three it withholds, and one
sentence Escoffier changed between two printings. If the record says
something lines 2439 to 2446 do not, or reads one of the three choices
above differently than you would,
&lt;a href="https://github.com/Alberto-Codes/saucier/issues/new?template=extraction.yml"&gt;the issue template&lt;/a&gt;
asks for the entry number and the source lines. Line 2440 is where I would
start.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;The release:&lt;/strong&gt; &lt;a href="https://github.com/Alberto-Codes/saucier/releases/tag/v0.6.0"&gt;saucier v0.6.0&lt;/a&gt;
— the first recorded procedure: &lt;code&gt;saucier show mornay&lt;/code&gt; prints six operations in
the book's own words, five quantities the text gives and three it withholds
left unresolved rather than guessed, every word checked against the source
before it prints. MIT.&lt;/p&gt;
</content:encoded>
    </item>
    <item>
      <title>I cut the last sauce off the file. Every line still parsed.</title>
      <link>https://alberto.codes/blog/2026-09-04-i-cut-the-last-sauce-off-the-file</link>
      <guid isPermaLink="true">https://alberto.codes/blog/2026-09-04-i-cut-the-last-sauce-off-the-file</guid>
      <pubDate>Fri, 04 Sep 2026 00:00:00 GMT</pubDate>
      <description>saucier now writes both copies of Escoffier as one stream, one record per line, so a shell can read the catalogue without importing my package. The first reader I wrote for it accepted an empty file, then a file missing its last sauce, with every remaining line valid JSON. Then a jq one-liner over the stream found a Périgueux the scan resolved and the text refused.</description>
      <content:encoded>&lt;p&gt;I deleted the last line of a 293-line file and fed the rest to a reader I
had written for it that week. It rebuilt two catalogues, printed a census one
sauce short, and exited zero. Every line it read was valid JSON. The line I
had deleted was the scan's Strawberry Sauce, entry 2417 of the 1907 text, at
line 41807, and nothing in the stream knew it was gone.&lt;/p&gt;
&lt;p&gt;The file exists because, after
&lt;a href="https://alberto.codes/blog/2026-09-03-the-book-spells-it-out-at-line-1437"&gt;the line-1437 post&lt;/a&gt;,
I wanted one list: every sauce that sits on half glaze, in both copies of the
book, side by side, asked from a shell and not from inside my own package. I
could not ask. The catalogue was two JSON files, one for each copy of the
book the project reads, and their shape was whatever my dataclasses were on
the day they were written. So &lt;code&gt;v0.5.0&lt;/code&gt; adds two commands. &lt;code&gt;saucier export&lt;/code&gt;
prints both catalogues to standard output as one stream, one complete JSON
record per line. &lt;code&gt;saucier import --check&lt;/code&gt; reads that stream back, rebuilds
every catalogue in memory, prints the census, and writes nothing. The JSON
files stay as they were, and &lt;code&gt;parse&lt;/code&gt;, &lt;code&gt;show&lt;/code&gt;, &lt;code&gt;tree&lt;/code&gt;, and &lt;code&gt;diff&lt;/code&gt; do not know
the stream exists.&lt;/p&gt;
&lt;h2&gt;One line, one record&lt;/h2&gt;
&lt;p&gt;Line 9 of the export is entry 25 of the 1909 text, the sauce Escoffier prints
at line 1467:&lt;/p&gt;
&lt;p&gt;&lt;img src="https://alberto.codes/saucier-one-record.svg" alt="One JSON record laid out with callouts. The envelope: schema saucier/1, type preparation, id escoffier-1909:line:1467, which is the catalogue and the line the heading sits on. Then catalogue escoffier-1909, title ORDINARY VELOUTÉ SAUCE, and a terms list with one entry: surface ORDINARY VELOUTÉ SAUCE, language fr, concept ordinary-veloute-sauce. Then parent pale-roux, with a note that null here means unresolved and never means the sauce has none. Then ref: entry 25, line 1467, fidelity transcription, the address a reader checks by hand. A body field follows, elided. Beneath: 293 lines, 2 catalogue records then 291 preparation records, 271,730 bytes, the same SHA-256 on every run." /&gt;&lt;/p&gt;
&lt;p&gt;Every line says what schema it follows, what kind of record it is, and which
catalogue it belongs to, so a line can be read on its own. The two catalogue
records come first, and each states how many preparations follow it.&lt;/p&gt;
&lt;h2&gt;Every line still parsed&lt;/h2&gt;
&lt;p&gt;Delete the last line of the export and feed the rest to the reader:&lt;/p&gt;
&lt;pre&gt;&lt;code class="language-console"&gt;$ uv run saucier export | head -n -1 | uv run saucier import --check
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;The first reader I wrote rebuilt two catalogues, printed the census with the
scan at 139 sauces instead of 140, and exited zero. Every line that survived
was complete JSON, and my reader asked nothing of the stream except that
each line parse.&lt;/p&gt;
&lt;p&gt;&lt;img src="https://alberto.codes/saucier-cut-stream.svg" alt="A column of 293 numbered lines. Line 1 is the catalogue record for escoffier-1909, stating 151 preparations; line 2 is the catalogue record for escoffier-1907, stating 140. Lines 3 to 292 are preparation records, drawn solid green. Line 293, STRAWBERRY SAUCE of the scan at line 41807, is drawn in dashed orange and struck through, labelled deleted. A note beside the green lines reads: every remaining line is valid JSON. An arrow from line 2 to the reader's verdict reads: line 2: catalogue 'escoffier-1907' states 140 preparations, the stream carries 139, exit 2. Caption: syntax cannot see a missing line; the record that promised a count can." /&gt;&lt;/p&gt;
&lt;p&gt;That is why a catalogue record states how many preparation records belong to
it. The reader counts them in, and the line number in its refusal points at
line 2, the record that made the promise, not at the gap. The count was not
in the stream until a review deleted a line and the reader said nothing. It
was the second of three things a stream of well-formed records could lie
about and the first reader could not see.&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;what the stream did&lt;/th&gt;
&lt;th&gt;the first reader&lt;/th&gt;
&lt;th&gt;the reader at &lt;code&gt;v0.5.0&lt;/code&gt;&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;arrived empty&lt;/td&gt;
&lt;td&gt;rebuilt nothing, exit 0&lt;/td&gt;
&lt;td&gt;&lt;code&gt;interchange carries no catalogues&lt;/code&gt;, exit 2&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;lost its last line&lt;/td&gt;
&lt;td&gt;census one short, exit 0&lt;/td&gt;
&lt;td&gt;&lt;code&gt;line 2: ... states 140 preparations, the stream carries 139&lt;/code&gt;, exit 2&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;repeated &lt;code&gt;parent&lt;/code&gt; on line 9&lt;/td&gt;
&lt;td&gt;Velouté moved to unresolved, exit 0&lt;/td&gt;
&lt;td&gt;&lt;code&gt;line 9: object repeats a key: ['parent']&lt;/code&gt;, exit 2&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;The empty stream matters because of the pipeline the command exists for,
&lt;code&gt;saucier export | saucier import --check&lt;/code&gt;. Without &lt;code&gt;pipefail&lt;/code&gt; a pipeline
reports its last command's exit status. If the export died before writing a
byte, the first reader saw an empty stream, accepted it, and the pipeline
would have finished with exit zero and the failed export behind it.&lt;/p&gt;
&lt;p&gt;The repeated key is a property of Python's JSON parser, which keeps the last
value of a duplicated key and says nothing. Append a second &lt;code&gt;&amp;quot;parent&amp;quot;:null&lt;/code&gt;
to line 9 and an ordinary decoder moves Velouté from derived to unresolved
with the census off by one and nothing raised. The reader still uses the
standard parser but gives it a hook that watches the keys arrive and refuses
a repeat. Nobody found that one. I went looking once the first two had taught
me to.&lt;/p&gt;
&lt;h2&gt;The one-liner, and the Périgueux row&lt;/h2&gt;
&lt;p&gt;Now the list I wanted, in one line of &lt;code&gt;jq&lt;/code&gt; that imports nothing of mine:&lt;/p&gt;
&lt;pre&gt;&lt;code class="language-console"&gt;$ uv run saucier export \
  | jq -r 'select(.type == &amp;quot;preparation&amp;quot; and .parent == &amp;quot;half-glaze&amp;quot;)
           | &amp;quot;\(.catalogue)  line \(.ref.line)  \(.title)&amp;quot;'
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Arranged into two columns:&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;escoffier-1909, the text&lt;/th&gt;
&lt;th&gt;escoffier-1907, the scan&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;line 1680  SAUCE BORDELAISE&lt;/td&gt;
&lt;td&gt;line 2057  SAUCE BORDELAISE&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;line 1708  BROWN CHAUD-FROID SAUCE&lt;/td&gt;
&lt;td&gt;line 2085  BROWN CHAUD=FROID SAUCE&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;line 1750  DEVILLED SAUCE&lt;/td&gt;
&lt;td&gt;line 2132  DEVILLED SAUCE&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;line 1841  ITALIAN SAUCE&lt;/td&gt;
&lt;td&gt;line 2233  ITALIAN SAUCE&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;line 1874  LYONNAISE SAUCE&lt;/td&gt;
&lt;td&gt;line 2268  LYONNAISE SAUCE&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;line 1886  MADEIRA SAUCE&lt;/td&gt;
&lt;td&gt;line 2280  MADEIRA SAUCE&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;line 2314  PERIQUEUX SAUCE&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;line 1930  PIQUANTE SAUCE&lt;/td&gt;
&lt;td&gt;line 2327  PIQUANTE SAUCE&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;line 1994  ROBERT SAUCE&lt;/td&gt;
&lt;td&gt;line 2398  ROBERT SAUCE&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;Eight sauces in the text, nine in the scan, and two rows that are findings.
&lt;code&gt;BROWN CHAUD=FROID&lt;/code&gt; is the scanner reading a hyphen as an equals sign at line
2085, and it stays as recorded:
&lt;a href="https://github.com/Alberto-Codes/saucier/blob/v0.3.0/docs/adr/0013-repair-structure-never-content.md"&gt;ADR-0013&lt;/a&gt;
repairs the punctuation that delimits a record, never the characters that
constitute one, and &lt;code&gt;=&lt;/code&gt; is inside the title.&lt;/p&gt;
&lt;p&gt;The unpaired row is the one I care about. In the 1909 text Périgueux refuses
to resolve, because its opening paragraph names two catalogued sauces, half
glaze and Madeira, and
&lt;a href="https://github.com/Alberto-Codes/saucier/blob/v0.3.0/docs/adr/0012-a-resolver-may-refuse-never-rank.md"&gt;the resolver may refuse but never rank&lt;/a&gt;.
The scan records half glaze, confidently. Same sentence, same two names, and
a page break between them:&lt;/p&gt;
&lt;p&gt;&lt;img src="https://alberto.codes/saucier-perigueux-page-break.svg" alt="Two panels of the same paragraph, entry 47, Périgueux Sauce. Left, escoffier-1909, the text: heading at line 1920, then one unbroken paragraph in which half-glaze at line 1923 and Madeira Sauce at line 1926 are both highlighted; the resolver reads two catalogued names and the verdict is parent unresolved, stated half-glaze, madeira-sauce. Right, escoffier-1907, the scan: heading at line 2314, half-glaze at line 2317 highlighted, then a blank line 2319 and the running page header 30 GUIDE TO MODERN COOKERY at line 2321, and only beyond them Madeira Sauce at line 2324, greyed out. The resolver reads the opening paragraph only, which ends at the blank line, so it sees one name and the verdict is parent half-glaze. Caption: the scan's paragraph ends at the blank line on 2319; the resolver never reads line 2324." /&gt;&lt;/p&gt;
&lt;pre&gt;&lt;code class="language-console"&gt;$ uv run saucier show perigueux-sauce
PÉRIGUEUX SAUCE
entry 47, line 1920, transcription of escoffier-1909
  term  PÉRIGUEUX SAUCE  [fr]  perigueux-sauce
  parent  (unresolved)
  stated  half-glaze, madeira-sauce
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;The resolver reads the opening paragraph only. In the scan that paragraph
ends at the blank line on 2319, the running header sits on 2321, and &amp;quot;per
quart of Madeira Sauce&amp;quot; is on 2324, outside what the resolver reads. One
name left, so it answers. That is the Aurore shape from
&lt;a href="https://alberto.codes/blog/2026-09-01-i-added-a-second-copy-of-the-same-book"&gt;the second-copy post&lt;/a&gt;,
where a damaged witness resolves what a clean one honestly cannot, and the
line-1437 post already counted Périgueux among three rows of it. What is new
is where I was standing when I saw it: in a shell, with a stream and &lt;code&gt;jq&lt;/code&gt;.
&lt;code&gt;grep -c '&amp;quot;parent&amp;quot;:null'&lt;/code&gt; on the same stream says 184, which is the 94
unresolved sauces of the text and the 90 of the scan.&lt;/p&gt;
&lt;h2&gt;What I am not claiming&lt;/h2&gt;
&lt;p&gt;This stream is not a database. It indexes nothing, keeps no history, and
answers no question about the graph of sauces. The JSON snapshot is still
the working store behind every other command, and the stream carries records
between processes and stops there. It is also less of a stream than the word
suggests: the reader consumes one line at a time, but a catalogue is
validated whole, so rebuilding one holds all of its preparations in memory
first.&lt;/p&gt;
&lt;p&gt;Two exports cannot be concatenated, because each carries every configured
catalogue and the reader stops at line 294 pointing back to line 1. A
catalogue id is a source id, which names a work and an edition, so two scans
of one printing would collide on every id. &lt;code&gt;saucier/1&lt;/code&gt; carries one catalogue
per source id and does not pretend the identifiers can tell those texts
apart.
&lt;a href="https://github.com/Alberto-Codes/saucier/blob/v0.5.0/docs/adr/0016-jsonl-is-the-interchange-not-a-store.md"&gt;ADR-0016&lt;/a&gt;
records those limits and why the interchange is not the store.&lt;/p&gt;
&lt;p&gt;And there is nothing here about what Escoffier changed between 1907 and
1909. The Périgueux row is a fact about a page header and a resolver that
reads one paragraph, surfaced from a shell.&lt;/p&gt;
&lt;p&gt;Everything here reproduces from the tag:&lt;/p&gt;
&lt;pre&gt;&lt;code class="language-console"&gt;$ git clone https://github.com/Alberto-Codes/saucier
$ cd saucier &amp;amp;&amp;amp; git checkout v0.5.0
$ uv sync &amp;amp;&amp;amp; uv run saucier parse
$ uv run saucier export | uv run saucier import --check
$ uv run saucier export | head -n -1 | uv run saucier import --check
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;The first pipeline prints 151 sauces and 140, then &lt;code&gt;2 catalogues and 291 preparations rebuilt. Nothing written.&lt;/code&gt; The second stops at line 2. Between
them is a 293-line file that anyone with &lt;code&gt;jq&lt;/code&gt; can ask about half glaze, and
that lists a Périgueux the scan resolved and the text refused. If you can
find a line in that stream that lies about the book and the reader lets
through,
&lt;a href="https://github.com/Alberto-Codes/saucier/issues/new?template=extraction.yml"&gt;the issue template&lt;/a&gt;
asks for the entry number and the source lines. Line 2314 of the scan is
where I would start.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;The release:&lt;/strong&gt; &lt;a href="https://github.com/Alberto-Codes/saucier/releases/tag/v0.5.0"&gt;saucier v0.5.0&lt;/a&gt;
— the stream and the reader that refuses it: &lt;code&gt;saucier export&lt;/code&gt; writes both
witnesses as &lt;code&gt;saucier/1&lt;/code&gt; records, 293 lines to the same SHA-256 every run, and
&lt;code&gt;saucier import --check&lt;/code&gt; rejects the empty file, the truncated one, the
repeated key and the rest, each with the line number where there is a line to
name — the empty stream has none. MIT.&lt;/p&gt;
</content:encoded>
    </item>
    <item>
      <title>The book spells it out at line 1437. Two of my posts said it never did.</title>
      <link>https://alberto.codes/blog/2026-09-03-the-book-spells-it-out-at-line-1437</link>
      <guid isPermaLink="true">https://alberto.codes/blog/2026-09-03-the-book-spells-it-out-at-line-1437</guid>
      <pubDate>Thu, 03 Sep 2026 00:00:00 GMT</pubDate>
      <description>Two earlier posts said Escoffier never spells out that half glaze is a reduction of Espagnole. Entry 23 does, in its first sentence. This is the clause of my own that hid it, the 27 entries that enter when it goes, and the ten derivations the release gives back on purpose.</description>
      <content:encoded>&lt;p&gt;Two earlier posts in this series carry a figure of the same four sauces:
&lt;a href="https://alberto.codes/blog/2026-08-19-there-is-no-model-in-this-parser"&gt;the one where ice cream got into the catalogue&lt;/a&gt;,
and &lt;a href="https://alberto.codes/blog/2026-08-21-marrow-sauce-finally-has-a-parent"&gt;the one where Marrow Sauce finally got a parent&lt;/a&gt;.
At the top of that figure is Espagnole. A dashed arrow runs from it down to
half glaze, and the arrow is labelled &lt;em&gt;assumed, not stated&lt;/em&gt; — half glaze, the
caption says, is &amp;quot;a reduction of Espagnole the book never spells out.&amp;quot; The
Marrow Sauce post goes further in prose: a &lt;em&gt;demi-glace&lt;/em&gt; is an Espagnole
reduction, &amp;quot;which every reader of Escoffier knew and the book therefore never
says.&amp;quot;&lt;/p&gt;
&lt;p&gt;saucier is a parser that reads Escoffier's 1909 &lt;em&gt;Guide to Modern Cookery&lt;/em&gt;
and records only what the book states. Here is entry 23 of that text, at
line 1437, in the chapter Escoffier titles &lt;em&gt;The Leading Warm Sauces&lt;/em&gt;:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;HALF GLAZE.&lt;/strong&gt; This is the Espagnole sauce, having reached the limit of
perfection by final despumation. It is obtained by reducing one quart of
Espagnole and one quart of first-class brown stock until its volume is
reduced to nine-tenths of a quart.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;The book says it. It says it in the first sentence of a numbered entry, in the
sauce chapter, nine entries before Bordelaise. I published two posts and two
figures asserting the book never says it, because my parser could not see
entry 23, and I believed my parser over the book.&lt;/p&gt;
&lt;p&gt;&lt;img src="https://alberto.codes/saucier-stated-chain-line-1437.svg" alt="The chain from brown roux to Marrow Sauce, five sauces, every arrow solid and labelled stated. Brown roux, entry 19 line 1317, states no parent. Espagnole, entry 22 line 1392, opens with &amp;quot;one lb. of brown roux&amp;quot; and records brown-roux. Half glaze, entry 23 line 1437, opens with &amp;quot;This is the Espagnole sauce&amp;quot; and records espagnole. Sauce Bordelaise, entry 32 line 1680, opens with &amp;quot;one-half pint of half-glaze&amp;quot; and records half-glaze. Marrow Sauce, entry 45 line 1895, &amp;quot;only a variety of the Bordelaise&amp;quot;, records sauce-bordelaise. A side note says the two earlier figures drew the top two arrows dashed and labelled them assumed, not stated; the book states both, and the parser could not see entries 19 to 23 because a rule of mine required their headings to name a mother." /&gt;&lt;/p&gt;
&lt;p&gt;The parser could not see it because the rule that decides what counts as a
sauce had two tests, and the second could overrule the first. This release
deletes the second test. Here is what that does to the census at &lt;code&gt;v0.4.0&lt;/code&gt;:&lt;/p&gt;
&lt;pre&gt;&lt;code class="language-console"&gt;$ uv run saucier parse
escoffier-1909  New and Revised Edition, January 1909 (impression: January 1920)
                transcription of Project Gutenberg 71395
                mothers: bechamel, espagnole, hollandaise, tomato, veloute
                151 sauces, 57 derived, 94 unresolved
escoffier-1907  no edition stated, copyright 1907
                ocr of Internet Archive cu31924000610117
                mothers: bechamel, espagnole, hollandaise, tomato, veloute
                140 sauces, 50 derived, 90 unresolved
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;At &lt;code&gt;v0.3.0&lt;/code&gt; the 1909 line read 124, 50, 74. Twenty-seven sauces entered.
Unresolved rose by twenty, and every one of the twenty is something the
parser could not see before: an entry my rule hid, or an ambiguity that was
invisible while the entry was hidden. The honest number got worse because the
instrument got better. And derived rose by seven, which is the misleading
number, for the best reason in the release.&lt;/p&gt;
&lt;h2&gt;The clause that vetoed Escoffier&lt;/h2&gt;
&lt;p&gt;Escoffier opens the chapter at line 1246 by saying what his sauces are:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Warm sauces are of two kinds: the leading sauces, also called &amp;quot;mother
sauces,&amp;quot; and the small sauces, which are usually derived from the
first-named, and are generally only modified forms thereof.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;The ice cream post introduced
&lt;a href="https://github.com/Alberto-Codes/saucier/blob/v0.1.0/docs/adr/0007-the-source-classifies-its-own-contents.md"&gt;ADR-0007&lt;/a&gt;:
the source decides what counts as a sauce, because deciding for ourselves is
what put vanilla ice cream in the catalogue. Under that record an entry
enters on two kinds of evidence. Its heading says &amp;quot;sauce&amp;quot;. Or the source
filed it in a chapter Escoffier titles as sauces — &lt;em&gt;and its heading also
names one of the five mothers.&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;That second test has two clauses. The chapter clause reads the source's own
classification. The mother clause adds a test of mine on top of the reading,
and when the two disagree, mine wins. &lt;strong&gt;A second test on a classification the
source has already made is a veto, not a check.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;&lt;img src="https://alberto.codes/saucier-veto-gate.svg" alt="Two rows of a decision flow. At v0.3.0: a numbered entry; does the heading say sauce, yes goes into the catalogue; no, is it in a chapter titled as sauces; no goes out; yes reaches a third test drawn in dashed orange, does the heading name a mother; no sends 27 entries out, yes goes in, and the catalogue holds 124. At v0.4.0 the third test is gone: an entry in a chapter titled as sauces goes straight in, and the catalogue holds 151." /&gt;&lt;/p&gt;
&lt;p&gt;I wrote the mother clause to keep &lt;code&gt;LENTEN ESPAGNOLE&lt;/code&gt; and &lt;code&gt;VELOUTÉ DE VOLAILLE&lt;/code&gt;, two derivatives in the sauce chapters whose headings never use the
word. It kept them.
&lt;a href="https://github.com/Alberto-Codes/saucier/blob/v0.4.0/docs/adr/0015-the-chapter-decides.md"&gt;ADR-0015&lt;/a&gt;
counts what it cost, in the three sauce chapters of the 1909 text:&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th style="text-align:right"&gt;entries&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;numbered entries in the three sauce chapters&lt;/td&gt;
&lt;td style="text-align:right"&gt;139&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;heading lacks the word &amp;quot;sauce&amp;quot;&lt;/td&gt;
&lt;td style="text-align:right"&gt;29&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;admitted on the mother clause&lt;/td&gt;
&lt;td style="text-align:right"&gt;2&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;vetoed&lt;/td&gt;
&lt;td style="text-align:right"&gt;27&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;The 27 are the three roux, half glaze, two gravies, a lobster method that
Escoffier numbered as its own entry, whisked mayonnaise, various cullises, and
18 compound butters from the chapter titled &lt;em&gt;Cold Sauces and Compound
Butters&lt;/em&gt;. Every one of them is an entry the source classified. My rule
classified them again and lost.&lt;/p&gt;
&lt;p&gt;Here is what that looked like on screen, at &lt;code&gt;v0.3.0&lt;/code&gt;:&lt;/p&gt;
&lt;pre&gt;&lt;code class="language-console"&gt;$ uv run saucier show robert-sauce --chars 260
ROBERT SAUCE
entry 52, line 1994, transcription of escoffier-1909
  term  ROBERT SAUCE  [en]  robert-sauce
  parent  (unresolved)

Finely mince a large onion and put it into a stewpan with butter. Fry
the onion gently and without letting it acquire any colour. Dilute
with one-third pint of white wine, reduce the latter by one-third,
add one pint of half-glaze, and leave to simmer for twen...
[378 more characters, raise --chars to read them]
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;The record and the sentence that contradicts it, four lines apart, for three
releases. Eleven sauces name half glaze in their opening paragraph. Five name
a roux. The book writes the chain in full — brown roux, Espagnole, half
glaze, Robert — and the catalogue had dropped the middle link.&lt;/p&gt;
&lt;p&gt;The new rule is one sentence: &lt;strong&gt;an entry inside a sauce chapter qualifies on
the chapter, and an entry outside one qualifies on its heading alone.&lt;/strong&gt; The
function that decides admission went from five lines to one:&lt;/p&gt;
&lt;pre&gt;&lt;code class="language-python"&gt;return names_a_sauce(title) or in_sauce_chapter
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;The mothers are still read, for the catalogue and for resolution, and they
take no part in admission. Nothing is hand-excluded to make the result tidy:
&lt;code&gt;VARIOUS CULLISES&lt;/code&gt; is entry 144, numbered inside a sauce chapter, and it is in
the catalogue now. An admitted entry that reads oddly as a preparation is a
finding, not a special case in the parser.&lt;/p&gt;
&lt;h2&gt;Twelve gained, and one of them was the bar&lt;/h2&gt;
&lt;p&gt;With half glaze and the roux in the catalogue, the chain resolver from the
Marrow Sauce post has something to resolve against. Twelve sauces now record
the parent Escoffier wrote, and none of them needed a new rule to do it:&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;resolves to&lt;/th&gt;
&lt;th&gt;sauces&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;half glaze&lt;/td&gt;
&lt;td&gt;Bordelaise, Brown Chaud-froid, Devilled, Italian, Lyonnaise, Madeira, Piquante, Robert&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;brown roux&lt;/td&gt;
&lt;td&gt;Espagnole&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;pale roux&lt;/td&gt;
&lt;td&gt;Velouté&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;white roux&lt;/td&gt;
&lt;td&gt;Scotch Egg&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;manied butter&lt;/td&gt;
&lt;td&gt;Mousseuse&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;Bordelaise is on that list, and the Marrow Sauce post made a point of
Bordelaise. Its opening names half a pint of half glaze and never says
Espagnole, so it stayed unresolved — &amp;quot;not a gap in this release,&amp;quot; I wrote,
but &amp;quot;the bar the model post will have to clear.&amp;quot; The bar is cleared, and no
model cleared it. A chapter test cleared it, by letting the resolver see an
entry that was in the book the whole time.&lt;/p&gt;
&lt;pre&gt;&lt;code class="language-console"&gt;$ uv run saucier tree espagnole
BROWN SAUCE OR ESPAGNOLE  [espagnole]  derives from brown-roux
├── HALF GLAZE  (en)
│   ├── SAUCE BORDELAISE  (fr)
│   │   └── MARROW SAUCE  (en)
│   ├── BROWN CHAUD-FROID SAUCE  (en)
│   ├── DEVILLED SAUCE  (en)
│   ├── ITALIAN SAUCE  (en)
│   ├── LYONNAISE SAUCE  (en)
│   ├── MADEIRA SAUCE  (en)
│   ├── PIQUANTE SAUCE  (en)
│   └── ROBERT SAUCE  (en)
├── LENTEN ESPAGNOLE  (fr)
│   └── GENEVOISE SAUCE  (en)
├── ORDINARY POIVRADE SAUCE  (en)
└── POIVRADE SAUCE FOR VENISON  (en)
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;That heading line is new. Escoffier names five mothers, and that does not
change. But Espagnole opens with &amp;quot;one lb. of brown roux dissolved in a tall,
thick saucepan with six quarts of brown stock,&amp;quot; and the catalogue can now see
the roux, so a mother may state a parent. It cannot see the stock: brown
stock is entry 7, in &lt;em&gt;Fonds de Cuisine&lt;/em&gt;, a chapter Escoffier does not title
as sauces. The chapter decides, and Chapter I decided stocks.&lt;/p&gt;
&lt;h2&gt;Ten lost, and they are the finding&lt;/h2&gt;
&lt;p&gt;Derived rose from 50 to 57, and seven is three numbers: twelve sauces gained
a parent, five admitted entries state one of their own, and &lt;strong&gt;ten sauces that
were resolved at &lt;code&gt;v0.3.0&lt;/code&gt; are unresolved now.&lt;/strong&gt; Eighteen of the 27 admissions
are compound butters, and seven of the ten losses are sauces that now see a
compound butter beside their old parent. The butters that came in are the
butters that took Cardinal's parent away.&lt;/p&gt;
&lt;p&gt;&lt;img src="https://alberto.codes/saucier-derived-waterfall.svg" alt="A waterfall chart. Derived at v0.3.0 is 50. Twelve sauces gained the parent Escoffier wrote, up to 62. Ten lost one to a butter or to half glaze, down to 52, drawn in dashed orange. Five admitted entries state a parent of their own, up to 57. Derived at v0.4.0 is 57. A note reads: net plus seven, read alone, is a gain that hides a loss." /&gt;&lt;/p&gt;
&lt;p&gt;Here is Cardinal, which is the shape of all ten, with a line that did not
exist at &lt;code&gt;v0.3.0&lt;/code&gt;:&lt;/p&gt;
&lt;pre&gt;&lt;code class="language-console"&gt;$ uv run saucier show cardinal-sauce
CARDINAL SAUCE
entry 69, line 2192, transcription of escoffier-1909
  term  CARDINAL SAUCE  [en]  cardinal-sauce
  parent  (unresolved)
  stated  bechamel, lobster-butter
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;And the paragraph beneath it, with the two verbs the resolver cannot read:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Boil&lt;/strong&gt; one pint of Béchamel, to which add one-half pint of fish &lt;em&gt;fumet&lt;/em&gt;
and a little truffle essence, and reduce by a quarter. &lt;strong&gt;Finish&lt;/strong&gt; the
sauce, when dishing up, with three tablespoonfuls of cream and three oz.
of very red lobster butter (No. 149).&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;At &lt;code&gt;v0.3.0&lt;/code&gt; Cardinal recorded &lt;code&gt;bechamel&lt;/code&gt;, and that was correct. Lobster
butter was on the page then too, at entry 149, but the mother clause had kept
it out of the catalogue, so as far as the resolver could see it was an
ingredient with no entry. Now it is catalogued, the opening paragraph states
two catalogued names, and
&lt;a href="https://github.com/Alberto-Codes/saucier/blob/v0.3.0/docs/adr/0012-a-resolver-may-refuse-never-rank.md"&gt;ADR-0012&lt;/a&gt;
says what happens then: the resolver refuses. It may not rank. So Cardinal is
unresolved, and so are Nantua, Noisette, Diplomate, Joinville, Herb and
Ravigote, each with at least one butter beside its old parent. Périgueux,
Reform and Chasseur lose theirs the same way to half glaze, which now sits
beside the base each of them names.&lt;/p&gt;
&lt;p&gt;The refusal is correct under the rule as written, and the rule as written is
wrong about this sentence. The source stated one &lt;em&gt;base&lt;/em&gt; and one &lt;em&gt;finish&lt;/em&gt;.
Any cook reading the paragraph knows which is which. The resolver reads
names. It cannot tell a base from a finish, because it has never been asked
to read the verb a name sits inside — and until it can, the honest answer is
the refusal.&lt;/p&gt;
&lt;p&gt;I could recover all ten this afternoon. A rule that prefers a mother when one
is stated would put Cardinal back on Béchamel and the derived count back near
67. I am not doing it, and the reason is
&lt;a href="https://alberto.codes/blog/2026-09-01-i-added-a-second-copy-of-the-same-book"&gt;the second-copy post's&lt;/a&gt;
Aurore sauce. The candidate rule is not tuned to save the number. Every
tuning of it that saves Cardinal is a choice made by me rather than a
statement made by the book, and the whole value of the unresolved count is
that it has never contained one of those.&lt;/p&gt;
&lt;p&gt;What the release does instead is make the refusal readable. That &lt;code&gt;stated&lt;/code&gt;
line names every catalogued candidate the opening paragraph states, in the
order the paragraph states them. The adjudication this tool cannot perform
still exists only in me. But the evidence I would adjudicate from is now on
screen next to the refusal, which is the most the parser is entitled to do.&lt;/p&gt;
&lt;p&gt;One of the twelve gains belongs in this section too. &lt;code&gt;MOUSSEUSE SAUCE&lt;/code&gt;
resolves to manied butter, because its opening calls for &amp;quot;one-half lb. of
stiffly-&lt;em&gt;manied&lt;/em&gt; butter&amp;quot; and manied butter is now a catalogued entry. The
same paragraph continues: &amp;quot;This preparation, though classified as a sauce,
is really a compound butter.&amp;quot; Escoffier classified it as a sauce and said in
the same breath that the classification was a courtesy. The catalogue reads
the chapter and records the sauce. The reader reads the sentence and knows
better. It stays as recorded. I would rather carry a visible oddity than an
invisible rule.&lt;/p&gt;
&lt;h2&gt;What it did to the scan&lt;/h2&gt;
&lt;p&gt;The 1907 witness moves from 115 sauces, 36 derived, 79 unresolved to 140, 50,
90. Twenty-five entries entered, not 27. &lt;code&gt;MONTPELLIER BUTTER&lt;/code&gt; is on the page
at line 3782 of the scan, and &lt;code&gt;HAZEL-NUT BUTTER&lt;/code&gt; at 3816:&lt;/p&gt;
&lt;pre&gt;&lt;code class="language-console"&gt;$ sed -n '3782p;3816p' corpus/escoffier-1907.txt
IS3— MONTPELLIER  BUTTER
15s— HAZEL-NUT  BUTTER
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;&lt;code&gt;IS3&lt;/code&gt; for 153, &lt;code&gt;15s&lt;/code&gt; for 155. The entry pattern wants digits and never
matches, so both sit inside the 284-entry blind spot the second-copy post
measured. Repairing them means deciding that &lt;code&gt;S&lt;/code&gt; is &lt;code&gt;5&lt;/code&gt; from outside the
document, and
&lt;a href="https://github.com/Alberto-Codes/saucier/blob/v0.3.0/docs/adr/0013-repair-structure-never-content.md"&gt;ADR-0013&lt;/a&gt;
makes no such decision.&lt;/p&gt;
&lt;p&gt;The scan also reads two entry numbers twice, 138 and 63, and until this
release the catalogue used the number as a preparation's identity, so the
later entry's bookkeeping overwrote the earlier one's. A preparation is now
identified by the line its heading sits on, which is unique in both witnesses
and is the field a reader checks by hand.&lt;/p&gt;
&lt;p&gt;The diff moves from 9 unmatched, 18 parent-changed, 35 ocr-suspected to 11,
19, 36. Three of the new rows are the Aurore shape again: &lt;code&gt;HERB&lt;/code&gt;, &lt;code&gt;RAVIGOTE&lt;/code&gt;
and &lt;code&gt;PÉRIGUEUX&lt;/code&gt; state two candidates in the proofread text and refuse, the
scan hides one candidate in each, and the scan answers. A damaged witness
confidently resolving what a clean witness honestly cannot was described once
as a finding. It has now happened four times.&lt;/p&gt;
&lt;h2&gt;Worse than the scanner, in one specific way&lt;/h2&gt;
&lt;p&gt;The second-copy post gave this series its second rule, &lt;em&gt;observed, never
assumed&lt;/em&gt;, and wrote it down as
&lt;a href="https://github.com/Alberto-Codes/saucier/blob/v0.3.0/docs/adr/0014-a-damaged-witness-cannot-establish-absence.md"&gt;ADR-0014&lt;/a&gt;:
a damaged witness cannot establish absence.&lt;/p&gt;
&lt;p&gt;I wrote that record about a scanner. Ten days later the same failure turned up
on the clean witness, in a proofread transcription with no OCR damage
anywhere in it. The blind spot was a clause I wrote, and the absence it
produced went into two posts, two figures, and the published claim that the
book never spells out what it spells out at line 1437. The project had 243
tests when I found it, and every one of them was green.&lt;/p&gt;
&lt;p&gt;The scan's blind spot could be measured, because there was a clean witness to
measure it against. The mother clause's blind spot had no second witness to
reveal it, because both copies of the book went through the same rule.&lt;/p&gt;
&lt;p&gt;&lt;img src="https://alberto.codes/saucier-two-books-one-reader.svg" alt="Two panels. Left, the scan's blind spot: the 1907 scan and the 1909 text both go into saucier diff, which measures 284 entries the scan cannot read, because a clean witness existed and the blind spot had a size. Right, the mother clause's blind spot: both books go into the same is_sauce rule, does the heading name a mother, and both catalogues come out missing the same 27 entries. Identical absences from two witnesses, and nothing left to compare against. Caption: a second copy of the book cannot catch an error in the reader; only a second reader can." /&gt;&lt;/p&gt;
&lt;p&gt;A second copy of the book cannot catch an error in the reader. Only a second
reader can, and the second reader here was a person following the chain
Robert states, by hand, and asking why the middle of it was not in the
catalogue.&lt;/p&gt;
&lt;p&gt;Both earlier posts now carry a dated note beside their figure. The figures
are left as printed, because what they show is genuinely what the parser
recorded at those tags. What they say about the book was wrong.&lt;/p&gt;
&lt;h2&gt;What I am not claiming&lt;/h2&gt;
&lt;p&gt;Not one of the ten lost derivations is a finding about Escoffier. Every one
is a finding about a resolver that reads names and not verbs. And the 27
admitted entries are a claim about where Escoffier put them, not about what a
sauce is — if a lobster method numbered as its own entry in the sauce chapter
is not a sauce, the argument is with the book's table of contents.&lt;/p&gt;
&lt;p&gt;Everything here reproduces from the tag:&lt;/p&gt;
&lt;pre&gt;&lt;code class="language-console"&gt;$ git clone https://github.com/Alberto-Codes/saucier
$ cd saucier &amp;amp;&amp;amp; git checkout v0.4.0
$ uv sync &amp;amp;&amp;amp; uv run saucier parse
$ uv run saucier show cardinal-sauce
$ uv run saucier tree espagnole
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;That prints 151 sauces and 140, the chain from brown roux to Robert that the
book has always written in full, and a &lt;code&gt;stated&lt;/code&gt; line under Cardinal naming the
two things it says, in the order it says them. If you can tell me a rule that
reads &amp;quot;finish the sauce with&amp;quot; as a finish and not a base, without reading
anything the sentence does not contain,
&lt;a href="https://github.com/Alberto-Codes/saucier/issues/new?template=extraction.yml"&gt;the issue template&lt;/a&gt;
asks for the entry number and the source lines. Entry 69, line 2192 is a
good place to start.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;The release:&lt;/strong&gt; &lt;a href="https://github.com/Alberto-Codes/saucier/releases/tag/v0.4.0"&gt;saucier v0.4.0&lt;/a&gt;
— the mother clause deleted: the second test is the chapter clause alone and
the heading test still stands beside it, the 27 entries
Escoffier filed among his own sauces come back, ten derivations are given up on
purpose, and a preparation is identified by the line its heading sits on rather
than by an entry number the scan reads twice. MIT.&lt;/p&gt;
</content:encoded>
    </item>
    <item>
      <title>Google's 4-bit Gemma already fit my 24 GiB card. I wanted the 20,000 tokens it left on the table.</title>
      <link>https://alberto.codes/blog/2026-09-02-googles-4-bit-gemma-already-fit-my-card</link>
      <guid isPermaLink="true">https://alberto.codes/blog/2026-09-02-googles-4-bit-gemma-already-fit-my-card</guid>
      <pubDate>Wed, 02 Sep 2026 00:00:00 GMT</pubDate>
      <description>The official Q4_0 build of Gemma 4 31B is 16.44 GiB and fits a 24 GiB card with room to spare, so "it fits" was never the claim. The budget is. I measured the decoder layer by layer, solved a 14.92 GiB pack that ties Google's build on four held-out benchmarks and wins one, and let the freed bytes buy context — 86,016 served tokens against 65,536, and 73,728 against 49,152 with an image aboard.</description>
      <content:encoded>&lt;p&gt;&lt;a href="https://huggingface.co/google/gemma-4-31B-it"&gt;Gemma 4 31B&lt;/a&gt; ships with something the last two models didn't have: an official quantization. Google publishes a &lt;a href="https://huggingface.co/google/gemma-4-31B-it-qat-q4_0-gguf"&gt;quantization-aware-trained Q4_0 GGUF&lt;/a&gt;, trained toward 4-bit deployment on purpose, at &lt;strong&gt;16.44 GiB&lt;/strong&gt;. On a 24 GiB card it loads with seven and a half gigabytes to spare.&lt;/p&gt;
&lt;p&gt;So for the first time in this series, &amp;quot;it fits&amp;quot; isn't a claim anyone needs to make. The question is what those spare gigabytes are worth — and on this model, they're worth more than usual. If you've been splitting a long transcript into pieces to fit a 24 GiB card, this one is for you. The only way to price the answer was to climb a ladder of context sizes until the card said no.&lt;/p&gt;
&lt;h2&gt;What a gigabyte of weights buys here&lt;/h2&gt;
&lt;p&gt;Everything a card holds past the weights goes to the KV cache, the per-token memory that grows with context. Gemma 4 31B's decoder has &lt;strong&gt;60 layers, and 50 of them are sliding-window&lt;/strong&gt;: each one sees the last 1,024 tokens and no more, so once a sequence passes that window its cache stops growing. Only the &lt;strong&gt;10 global layers&lt;/strong&gt; keep growing with context — on a runtime that honors the sliding window, as llama.cpp does. One that allocates full-context cache for every layer would price this model six times higher, and the trade below would vanish.&lt;/p&gt;
&lt;p&gt;That makes the price of context small and flat. Measured at the runtime — llama.cpp b10362, not the config file — a token of context costs &lt;strong&gt;81,920 bytes&lt;/strong&gt; once the windows saturate, plus a fixed 1,200 MiB pool per sequence for the fifty windowed layers. (The config's arithmetic says half that per token, because Gemma stores one tensor for both keys and values; the runtime allocates both anyway, and the runtime is what runs.) A gigabyte of weights freed is about &lt;strong&gt;13,100 tokens&lt;/strong&gt; of context bought.&lt;/p&gt;
&lt;p&gt;&lt;img src="https://alberto.codes/vramfit-24gib-kv-geometry.svg" alt="A 24 GiB card, twice. Google's QAT Q4_0 holds 16.44 GiB of weights and serves 65,536 tokens of text context. This pack holds 14.92 GiB and serves 86,016. The gap in weights is 1.5 GiB; the gap in served context is 20,480 tokens. Both are measured load boundaries on an RTX 4090, llama.cpp b10362, one sequence, 4,096-token rungs, measured 2026-08-31." /&gt;&lt;/p&gt;
&lt;h2&gt;The budget is the claim&lt;/h2&gt;
&lt;p&gt;That inverts how &lt;a href="https://github.com/Alberto-Codes/vramfit"&gt;vramfit&lt;/a&gt; had been thinking. Publications &lt;a href="https://alberto.codes/blog/2026-08-11-i-couldnt-tell-my-quantized-model-from-the-baseline"&gt;#1&lt;/a&gt; and &lt;a href="https://alberto.codes/blog/2026-08-22-the-2-bit-label-was-4-5-bits-inside"&gt;#2&lt;/a&gt; asked how few bytes a model could survive. This one asks the other direction: pick the context you want, subtract, and give the weights whatever's left. Nine gigabytes of KV headroom on a 24 GiB card leaves 15 GiB for weights — call that arm kv9. That's the one that shipped. An eleven-gigabyte headroom leaves 13 GiB, and that arm is the more interesting story, because it lost.&lt;/p&gt;
&lt;h2&gt;Measuring a model that only speaks in turns&lt;/h2&gt;
&lt;p&gt;Before either arm could be solved, the meter had to be fixed. This checkpoint is instruction-tuned and quantization-aware-trained, and it is &lt;em&gt;channel-locked&lt;/em&gt;: it has only ever seen text inside its chat frame. Feed it raw prose and it prices it as noise. On the calibration corpus, the same tool read a perplexity in the thousands on bare text and &lt;strong&gt;37.39&lt;/strong&gt; on the identical text wrapped in the model's turn markers.&lt;/p&gt;
&lt;p&gt;A sensitivity map measured on the raw distribution would price every layer against text no user ever sends. So every measurement in this campaign — the scan, the importance matrix, the perplexity and divergence readings — ran inside one fixed model-turn frame: 357 blocks of public-domain text, 182,404 tokens, the turn markers parsed as the control tokens they are. One trap that step caught: the bf16 conversion had emitted &lt;code&gt;&amp;lt;|turn&amp;gt;&lt;/code&gt; as an ordinary token, and only a token count exposed it.&lt;/p&gt;
&lt;p&gt;The map itself is keyed per decoder layer, 60 layers and the token embedding, because this is a dense model and a layer is the unit the pack addresses. The vision tower never entered it. A text-measured map licenses no vision claim, and that gap gets its own measurement below.&lt;/p&gt;
&lt;h2&gt;The recipe: cheap early, protected deep&lt;/h2&gt;
&lt;p&gt;The solver walked the map greedily by damage per byte saved, under the 15 GiB weight budget, over Google's &lt;a href="https://huggingface.co/google/gemma-4-31B-it-qat-q4_0-unquantized"&gt;QAT-unquantized checkpoint&lt;/a&gt; — the weights Google annealed toward 4-bit before quantizing them. What came out:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;6 layers at &lt;code&gt;Q2_K&lt;/code&gt;&lt;/strong&gt; — layers 1 through 5, and 11. The map priced the early layers cheapest.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;9 layers at &lt;code&gt;Q3_K&lt;/code&gt;&lt;/strong&gt; — layer 0, 6 through 10, and 12 through 14.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;45 layers at &lt;code&gt;Q4_K&lt;/code&gt;&lt;/strong&gt; — layers 15 through 59. The solve spent its budget protecting depth.&lt;/li&gt;
&lt;li&gt;The token embedding and the output head at &lt;code&gt;Q4_K&lt;/code&gt;.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;img src="https://alberto.codes/vramfit-24gib-recipe.svg" alt="Sixty decoder layers in a strip. Layers 1 to 5 and 11 sit at Q2_K, layer 0, 6 to 10, and 12 to 14 at Q3_K, and layers 15 to 59 at Q4_K. Beside the strip, the two shipped files: the 14.92 GiB decoder and the 629 MiB projector sidecar." /&gt;&lt;/p&gt;
&lt;p&gt;The packed decoder is &lt;strong&gt;14.92 GiB&lt;/strong&gt;, 86.08 MiB under budget, with the recipe's 81-step trace beside it. Google's build is one preset applied uniformly; this one buys forty-five protected layers by spending fifteen cheap ones, and still lands a gigabyte and a half lighter.&lt;/p&gt;
&lt;p&gt;Images need a second file: the &lt;em&gt;projector&lt;/em&gt;, the encoder that turns pixels into tokens the decoder can read. It ships beside the decoder as a sidecar, priced below.&lt;/p&gt;
&lt;h2&gt;The metric that pointed backwards&lt;/h2&gt;
&lt;p&gt;The 13 GiB arm, kv11, was the aggressive point, and on the lingua-franca metric it &lt;em&gt;won&lt;/em&gt;. Perplexity ratio against the bf16 reference, in frame:&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Arm&lt;/th&gt;
&lt;th&gt;Weights&lt;/th&gt;
&lt;th&gt;PPL / bf16 ↓&lt;/th&gt;
&lt;th&gt;Mean KLD ↓&lt;/th&gt;
&lt;th&gt;Same top ↑&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;kv11 (13 GiB)&lt;/td&gt;
&lt;td&gt;12.92 GiB&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;1.0390&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;0.1352&lt;/td&gt;
&lt;td&gt;86.35%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;kv9 (15 GiB, shipped)&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;14.92 GiB&lt;/td&gt;
&lt;td&gt;1.0681&lt;/td&gt;
&lt;td&gt;0.0446&lt;/td&gt;
&lt;td&gt;92.04%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Google's QAT Q4_0&lt;/td&gt;
&lt;td&gt;16.44 GiB&lt;/td&gt;
&lt;td&gt;1.1043&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;0.0420&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;92.32%&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;Perplexity ratio ranks the three packs in one order; every fidelity metric ranks them in the &lt;em&gt;opposite&lt;/em&gt; order. kv11 has the best perplexity ratio and &lt;strong&gt;3.2 times the divergence&lt;/strong&gt; of Google's build. A perplexity-only scoreboard would have shipped the wrong file. I don't know why perplexity pointed backwards on this model, and this post doesn't pretend to.&lt;/p&gt;
&lt;p&gt;&lt;img src="https://alberto.codes/vramfit-24gib-two-arms.svg" alt="Three packs, two rankings. Perplexity ratio orders kv11, kv9, then QAT. Mean KL divergence orders QAT, kv9, then kv11. The two orderings are exactly reversed, and the held-out benchmarks below break the tie." /&gt;&lt;/p&gt;
&lt;p&gt;The held-out benchmarks sided with divergence. Five tasks fixed before any run, full evaluation splits, every arm on the same lane. A delta inside the combined standard error is a tie:&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Task&lt;/th&gt;
&lt;th&gt;kv9 (shipped)&lt;/th&gt;
&lt;th&gt;Google's QAT Q4_0&lt;/th&gt;
&lt;th&gt;Verdict&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;MMLU (5-shot)&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;71.36&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;70.20&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;win (+1.15, σ 0.53)&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GSM8K (5-shot)&lt;/td&gt;
&lt;td&gt;92.34&lt;/td&gt;
&lt;td&gt;92.42&lt;/td&gt;
&lt;td&gt;tie&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;HellaSwag (10-shot)&lt;/td&gt;
&lt;td&gt;58.71&lt;/td&gt;
&lt;td&gt;59.34&lt;/td&gt;
&lt;td&gt;tie (−0.63, σ 0.69)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Winogrande (5-shot)&lt;/td&gt;
&lt;td&gt;68.27&lt;/td&gt;
&lt;td&gt;68.03&lt;/td&gt;
&lt;td&gt;tie&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;ARC-Challenge (25-shot)&lt;/td&gt;
&lt;td&gt;61.77&lt;/td&gt;
&lt;td&gt;61.09&lt;/td&gt;
&lt;td&gt;tie&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;Four ties and a win, from the lighter file. kv11 took the same slice and lost HellaSwag outright, &lt;strong&gt;−2.33 points against a combined σ of 0.69&lt;/strong&gt;, with four ties around it. Its divergence showed up on exactly one task, which is one more than the shipping bar allows. kv11 would have served the most context of the three, and it stays &lt;a href="https://github.com/Alberto-Codes/vramfit/issues/423"&gt;on the record&lt;/a&gt; as the arm that measured too much damage to publish.&lt;/p&gt;
&lt;p&gt;Two honest asymmetries travel with the first table. Google's build holds a slightly better mean KL divergence and top-token agreement than the shipped pack, and the benchmarks are what settle that disagreement. And that table is measured on the pack's own calibration frame — the same corpus its importance matrix consumed — which leans the in-frame numbers toward my pack. The held-out slice is the check on that.&lt;/p&gt;
&lt;h2&gt;What the bytes bought, served&lt;/h2&gt;
&lt;p&gt;Computed capacity is arithmetic. Served capacity is the ladder: load the file at a context size, step up 4,096 tokens, repeat until the load fails. Both packs ran the ladder on the same RTX 4090, the same day, llama.cpp b10362 Vulkan, one sequence:&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Serving shape&lt;/th&gt;
&lt;th&gt;This pack&lt;/th&gt;
&lt;th&gt;Google's QAT Q4_0&lt;/th&gt;
&lt;th&gt;Gain&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Text only&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;86,016&lt;/strong&gt; tokens&lt;/td&gt;
&lt;td&gt;65,536&lt;/td&gt;
&lt;td&gt;+20,480 (+31.25%)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;One image aboard&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;73,728&lt;/strong&gt; tokens&lt;/td&gt;
&lt;td&gt;49,152&lt;/td&gt;
&lt;td&gt;+24,576 (+50%)&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;The image row moves two variables at once — my decoder &lt;em&gt;and&lt;/em&gt; my converted projector — so the card also prints the one-variable version: behind Google's own BF16 projector, this pack serves an image at 69,632 tokens, still +41.7%. The projector conversion adds the last 4,096.&lt;/p&gt;
&lt;p&gt;&lt;code&gt;fit24gib&lt;/code&gt; is a contract, not a boast, and the boundary has edges:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;It moves with the frame&lt;/strong&gt; — in this project's vocabulary, the exact conditions a number was measured under: the box, the build, the day, the idle VRAM. The 2026-08-28 ladder found 81,920 for text; three days later, in a different frame, the same file passed 86,016. The boundary moves with the box's idle VRAM share. Both are real, and the card names both.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Pass &lt;code&gt;-np 1&lt;/code&gt;.&lt;/strong&gt; The server defaults to four slots, which adds about 2,400 MiB of sliding-window cache on this geometry and fails loads that fit at one.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Keep about 200 MiB free when serving images&lt;/strong&gt;, and cap the encode batch at one image. The image-encode buffer allocates at request time, and this server build crashes on that failure instead of refusing.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;The ladder is a fit bar, not a speed bar.&lt;/strong&gt; The boundary check decoded five tokens.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Throughput at the boundary is unmeasured, but throughput at a working context is not, and it's the number I refused to print last time. Both quantized arms ran the same 20-task subset at 8,192 tokens of context on the 4090, the target card: &lt;strong&gt;47.8 tokens per second for this pack against 43.3&lt;/strong&gt; for Google's. On an H100 the order flipped, with Google's build 15% ahead. Twenty generations, one slot, my pack's answers averaging 20 tokens to Google's 48 — a measured number on the card the claim is about, with its caveats attached.&lt;/p&gt;
&lt;h2&gt;The vision claim had to be earned separately&lt;/h2&gt;
&lt;p&gt;The map measured text, and a card that infers image quality from text damage would be guessing. So the campaign measured it: a BF16 reference decoder generated greedily over ten held-out 768×768 images, and each quantized arm was teacher-forced along the reference's sequence, reading a truncated top-20 KL divergence at the server boundary (the server caps its probability list at 20). Positions that carry image content are scored apart from the chat-frame policy tokens that dominate the average.&lt;/p&gt;
&lt;p&gt;Content-class results, 120 positions:&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Arm&lt;/th&gt;
&lt;th&gt;Mean KLD ↓&lt;/th&gt;
&lt;th&gt;p95 ↓&lt;/th&gt;
&lt;th&gt;Same top ↑&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;This pack, as shipped&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;0.0050&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;0.0193&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;99.2%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;This pack, Google's BF16 projector&lt;/td&gt;
&lt;td&gt;0.0045&lt;/td&gt;
&lt;td&gt;0.0239&lt;/td&gt;
&lt;td&gt;99.2%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Google's QAT Q4_0, Google's BF16 projector&lt;/td&gt;
&lt;td&gt;0.0373&lt;/td&gt;
&lt;td&gt;0.1928&lt;/td&gt;
&lt;td&gt;97.5%&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;&lt;strong&gt;7.5 times below Google's build, 47 times above the instrument's noise floor&lt;/strong&gt; (1.07e-4, measured by scoring an arm against its own greedy output). The reference ran on CPU — a second instrument — so these read as divergence from the reference, never as same-instrument damage; only the pack-against-Q4_0 comparison shares one instrument.&lt;/p&gt;
&lt;p&gt;That measurement is also what priced the projector. Google's ships in BF16 at 1,145 MiB. Run through llama-quantize's &lt;code&gt;Q4_K_M&lt;/code&gt; recipe it lands at &lt;strong&gt;629 MiB&lt;/strong&gt; — and holds not one &lt;code&gt;Q4_K&lt;/code&gt; tensor, because every quantizable tensor in it falls back on this geometry to &lt;code&gt;Q5_0&lt;/code&gt; or &lt;code&gt;Q8_0&lt;/code&gt;. The recipe name labels the command, not the contents. The cost was 0.0045 to 0.0050 on the content mean; the return was 482 MiB at load and one rung of context. It shipped.&lt;/p&gt;
&lt;p&gt;Then a second campaign put the bound somewhere real: &lt;strong&gt;1,349 GUI screenshots&lt;/strong&gt; at 1280×720 from a public computer-use dataset, three arms on one H100. On image-content positions this pack diverges less than Google's build — 0.0973 against 0.1143 at the mean, and a median 3.6 times lower. Read that as a bound, not a separation: this campaign measured no noise floor, and a prompt-prefix difference between the three files rides inside the mean margin. On a task-identification score, no arm separates: the BF16 reference itself scored 49.9%, my pack 51.2%, Google's 51.4% (from a 512-token regeneration, after every one of its 48-token answers ended mid-thought), all inside a 6.9% judge noise floor. One screenshot against a whole task's name is a ceiling the reference can't clear either.&lt;/p&gt;
&lt;h2&gt;What I'm still not claiming&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;The sensitivity map and the importance matrix are not published.&lt;/strong&gt; Publications #1 and #2 shipped their maps as datasets; this one ships the recipe, the run log, the evaluation sidecars, and both campaign records, but the map and the matrix stay in the run archive. The recipe replays the type placement from the base checkpoint. Without the map, &lt;code&gt;vramfit plan&lt;/code&gt; cannot re-solve this model for a different budget, and without the matrix a rebuild reproduces the placement but not this file's exact bytes.&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Note 2026-09-04: both have since been published. The &lt;a href="https://huggingface.co/datasets/Alberto-Codes/gemma-4-31B-it-sensitivity-maps"&gt;sensitivity-map dataset&lt;/a&gt; carries the map, the framed importance matrix, and both calibration texts, so neither limit above still holds.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;&lt;strong&gt;Google's build is the comparator, not the shelf.&lt;/strong&gt; It is the vendor's own artifact and the obvious baseline. No claim here covers the other GGUFs of this model on the Hub.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;The vision numbers bound divergence, not safety.&lt;/strong&gt; Quantization compresses every tensor with one lossy procedure and can shift any behavior. The tables are the measured bound on that shift over text, ten images, and 1,349 screenshots. Read the card as a damage disclosure. Deploy the pack with whatever protections you'd give the base model.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Damage values don't travel.&lt;/strong&gt; The map's 0.184 predicted damage is one scan's number in one frame. It ranks layers within this map and nothing else.&lt;/p&gt;
&lt;h2&gt;Where it lives&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;The model:&lt;/strong&gt; &lt;a href="https://huggingface.co/Alberto-Codes/gemma-4-31B-it-fit24gib-GGUF"&gt;gemma-4-31B-it-fit24gib-GGUF&lt;/a&gt; — two files, one artifact. The decoder serves text alone; add the projector sidecar for images. The card carries the recipe, the serve ladders, the eval sidecars, both campaign records, and the reproduce commands. Apache 2.0, under the Gemma 4 license note.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;The tool:&lt;/strong&gt; &lt;a href="https://github.com/Alberto-Codes/vramfit/releases/tag/v0.4.0"&gt;vramfit v0.4.0&lt;/a&gt; — what this campaign forced: per-layer KV geometry priced at the runtime's measured allocation, a &lt;code&gt;capacity&lt;/code&gt; readout that turns a packed recipe back into tokens, the projector sidecar and vision line, and the framed-calibration script. MIT.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;The policy:&lt;/strong&gt; &lt;a href="https://github.com/Alberto-Codes/vramfit/blob/v0.4.0/docs/adr/0030-vision-budget-sidecar.md"&gt;ADR-0030&lt;/a&gt; — how a vision tower enters the budget, and what a text-measured map may claim.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;The how-to:&lt;/strong&gt; &lt;a href="https://alberto.codes/blog/2026-08-15-fit-a-model-to-the-gpu-you-actually-have"&gt;fit a model to the GPU you actually have&lt;/a&gt; — the command-by-command version.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Last time the villain was a label that hid a fallback. This time there was no villain. Google's build is good, it fits, and it's the first comparator in this series I'd happily run. The only thing wrong with it was the seven and a half gigabytes it wasn't using — and the only way to find out what they were worth was to measure, solve, and serve the ladder until it failed.&lt;/p&gt;
</content:encoded>
    </item>
    <item>
      <title>I added a second copy of the same book. It told me nothing about the book.</title>
      <link>https://alberto.codes/blog/2026-09-01-i-added-a-second-copy-of-the-same-book</link>
      <guid isPermaLink="true">https://alberto.codes/blog/2026-09-01-i-added-a-second-copy-of-the-same-book</guid>
      <pubDate>Tue, 01 Sep 2026 00:00:00 GMT</pubDate>
      <description>Post three of the saucier series puts a second witness of Escoffier into the corpus — the 1907 first printing, as a scan — and gets back no facts about cookery and four classes of fact about its own instruments. The number of sauces the revision supposedly added went twenty, then eight, then none. Along the way the project discovered it had been citing the wrong edition since its first commit.</description>
      <content:encoded>&lt;p&gt;I put a second copy of Escoffier into the corpus this week. The project had
been reading one text of &lt;em&gt;A Guide to Modern Cookery&lt;/em&gt; since its first commit,
and there is a second one available — the 1907 first printing, photographed
and machine-read by the Internet Archive. Two witnesses of one book, and a
command that compares them.&lt;/p&gt;
&lt;p&gt;Two things came out of it that are worth having. The catalogue now reads its
own edition out of the book's front matter, rather than taking it from the
name of a file on my disk — which is infrastructure the rest of this series
will stand on, and which did not exist a week ago. And the 1907 printing,
which the parser could not read at all when it arrived, now yields 115
sauces.&lt;/p&gt;
&lt;p&gt;And then the result, which is the reason for the post: &lt;strong&gt;it told me nothing
about the book.&lt;/strong&gt; Not one confirmed editorial difference between the two
printings. Every apparent difference I have found so far has turned out to
be a property of my instruments — a wrapped heading, a broken separator, a
corrupted digit, a gap in my own comparison code. The release produced no
facts about a cookbook and four classes of fact about the tools reading it.&lt;/p&gt;
&lt;p&gt;That is a better outcome than it sounds. The failure underneath it has a
shape worth naming before the cookbook details start, because it is not
about cookbooks: &lt;strong&gt;an instrument reported an absence it had no way to
observe, through a blind spot nobody had measured.&lt;/strong&gt; It said a book lacked
something the book contains. Then, having been corrected, it did it again,
and the second time the gates were just as green as the first.&lt;/p&gt;
&lt;p&gt;If you run anything that compares two versions of a document and reports what
changed, that failure is available to you today, and the question this post is
really asking is whether your pipeline could tell you how much of its input it
could not read.&lt;/p&gt;
&lt;h2&gt;First, the label was wrong&lt;/h2&gt;
&lt;p&gt;Before any of that: this project had been citing an edition it had never
parsed.&lt;/p&gt;
&lt;p&gt;The corpus file was called &lt;code&gt;escoffier-1907&lt;/code&gt;. It is not the 1907 edition. Here
is the printing history, transcribed in the file itself, forty lines above
the first recipe:&lt;/p&gt;
&lt;pre&gt;&lt;code class="language-console"&gt;$ sed -n '119,126p' corpus/escoffier-1909.txt
        _First Printed, May 1907
     Second Impression, December 1907
  New and Revised Edition, January 1909
 New Impressions, August 1911, May 1913,
        March 1916, January 1920._


_Copyright 1907 by William Heinemann._
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;The book is the New and Revised Edition of January 1909, in its January 1920
impression. The only 1907 on the page is a copyright line and the date of a
first printing this file is not. I named the file after the most
1907-looking thing in view, and thereafter the code read the filename.&lt;/p&gt;
&lt;p&gt;The claims survive this. Every entry number, every line number, every
derivation was correct and still reproduces. What was wrong was the label on
the book they cite — which, for a record whose whole pitch is &lt;em&gt;go and check&lt;/em&gt;,
is not a small field to get wrong. It is also
&lt;a href="https://github.com/Alberto-Codes/saucier/blob/v0.1.0/docs/adr/0007-the-source-classifies-its-own-contents.md"&gt;ADR-0007&lt;/a&gt;
broken by the person who wrote it. That rule says the source decides what
counts as a sauce, because deciding for ourselves is what put vanilla ice
cream in the catalogue in post #1. A book states its edition more plainly
than it states anything else. I read the filename anyway, in the same release
that introduced the rule.&lt;/p&gt;
&lt;p&gt;So a source now reports the edition it states, and the &lt;code&gt;source_id&lt;/code&gt; derives
from that reading. Four facts come out of the front matter and are kept
apart, because one string cannot carry four: the edition statement, its year,
the last impression, and the copyright year. With no edition stated, the
copyright year decides — a first printing has no printing history to print. A
source that states neither raises rather than guessing.&lt;/p&gt;
&lt;pre&gt;&lt;code class="language-json"&gt;{ &amp;quot;source_id&amp;quot;: &amp;quot;escoffier-1909&amp;quot;, &amp;quot;work&amp;quot;: &amp;quot;escoffier&amp;quot;,
  &amp;quot;origin&amp;quot;: &amp;quot;Project Gutenberg 71395&amp;quot;, &amp;quot;fidelity&amp;quot;: &amp;quot;transcription&amp;quot;,
  &amp;quot;edition&amp;quot;: { &amp;quot;statement&amp;quot;: &amp;quot;New and Revised Edition, January 1909&amp;quot;,
               &amp;quot;stated_year&amp;quot;: 1909, &amp;quot;impression&amp;quot;: &amp;quot;January 1920&amp;quot;,
               &amp;quot;copyright_year&amp;quot;: 1907 } }
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Which frees the name for the real 1907, and both published posts in this
series get a correction: they describe the book as Escoffier's 1907 &lt;em&gt;Guide to
Modern Cookery&lt;/em&gt;, and it is the 1909 revision. I have corrected the prose in
both, the way &lt;a href="https://alberto.codes/blog/2026-08-11-i-couldnt-tell-my-quantized-model-from-the-baseline"&gt;the August 11 post&lt;/a&gt;
was fixed twice — the canonical post is the one that gets corrected. Each also
carries a dated note beside its console block, because the identifier those
posts printed is genuinely what the code printed, and a reader running their
pinned commands will still see it. The counts, entries and line numbers are
unaffected and reproduce exactly as published.&lt;/p&gt;
&lt;h2&gt;Twenty, then eight, then none&lt;/h2&gt;
&lt;p&gt;Now the second witness, and the arc that is the actual story.&lt;/p&gt;
&lt;p&gt;&lt;code&gt;saucier diff&lt;/code&gt; compares two stored catalogues and puts a cause on every row
where they disagree. The first time it ran, it reported &lt;strong&gt;twenty sauces the
1909 revision had added&lt;/strong&gt; — twenty preparations present in the later
catalogue and absent from the earlier one.&lt;/p&gt;
&lt;p&gt;All twenty are printed in the 1907 book. I checked every one by hand. They
were invisible to the parser because the entry pattern is
&lt;code&gt;^(\d{1,4})—(.+)$&lt;/code&gt;, and a photographed page defeats it three ways: the digits
come through as letters, the em dash comes through as something else, or a
space lands in the middle.&lt;/p&gt;
&lt;p&gt;Six of them, quoted from &lt;code&gt;corpus/escoffier-1907.txt&lt;/code&gt; with their line numbers
and arranged into a table:&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;line&lt;/th&gt;
&lt;th&gt;as the scan has it&lt;/th&gt;
&lt;th&gt;line&lt;/th&gt;
&lt;th&gt;as the scan has it&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;3003&lt;/td&gt;
&lt;td&gt;&lt;code&gt;loi— POULETTE  SAUCE&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;2072&lt;/td&gt;
&lt;td&gt;&lt;code&gt;33- CHASSEUR  SAUCE  (Escoffier's  Method)&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;3147&lt;/td&gt;
&lt;td&gt;&lt;code&gt;Ill— WHITE  WINE  SAUCE&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;3345&lt;/td&gt;
&lt;td&gt;&lt;code&gt;126-- MAYONNAISE  SAUCE&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2560&lt;/td&gt;
&lt;td&gt;&lt;code&gt;6s— BERCY  SAUCE&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;3172&lt;/td&gt;
&lt;td&gt;&lt;code&gt;1 12- APPLE  SAUCE&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;&lt;code&gt;33- CHASSEUR SAUCE&lt;/code&gt; has perfectly legible digits and an ordinary hyphen. The
parser wanted an em dash, so the sauce did not exist, and the diff reported
that Escoffier had added it in 1909.&lt;/p&gt;
&lt;p&gt;Repairing those separators took two rounds and moved the 1907 census from 102
sauces to 113, then to 115. The twenty claimed additions fell to eight — plus
one claimed removal, which had the same cause pointing the other way.&lt;/p&gt;
&lt;p&gt;Then the third round, which is the one that matters. Eight was still a claim
of absence, and absence is not a thing this instrument can observe. So the
cause is gone. There is no &lt;code&gt;added&lt;/code&gt; row against a scanned witness any more:&lt;/p&gt;
&lt;pre&gt;&lt;code class="language-console"&gt;  9 unmatched, 18 parent-changed, 35 ocr-suspected
  entries read  2679 of escoffier-1907, 2963 of escoffier-1909, a blind spot of 284
  A witness is ocr. An unmatched row says the diff found no
  counterpart, never that the printing lacks one.
  No row is adjudicated. An ocr-suspected row is a suspicion, and
  separating a scan artefact from a revision needs both lines read.
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Twenty, then eight, then none. Three rounds of being confidently wrong about
the same question, on a project with 223 tests, 100% coverage, and fourteen
decision records — and every one of the three was caught by a person reading
output, never by a test. The gates were green the entire time. They were
green for the twenty, and they are green now.&lt;/p&gt;
&lt;p&gt;Two more of the same species turned up alongside, and they are worth a
sentence each because they came from opposite ends of the pipeline. An entry
heading that wrapped onto a second line lost its tail, and since the two
printings are set to different column widths, they wrapped in different
places — so the two witnesses recorded visibly different titles for a heading
Escoffier printed identically, and the diff blamed the scan for a difference
my own regular expression had manufactured. At the other end, the comparison
was intersecting its two indexes, so any preparation whose names had been
matched across the witnesses never had its parent compared at all. Six real
disagreements were invisible. Neither defect failed anything.&lt;/p&gt;
&lt;h2&gt;Aurore Sauce, and the failure worth being frightened of&lt;/h2&gt;
&lt;p&gt;Everything above is an instrument reporting something that is not there. Here
is the other direction, which is worse.&lt;/p&gt;
&lt;p&gt;Entry 60 is Aurore Sauce. Here is its opening, as the 1909 transcription
gives it:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Into one-half pint of boiling velouté put the same quantity of very red
tomato purée (No. 29), and mix the two.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;Two candidates are named there: velouté and tomato. Post #2 built its whole
argument on what the parser does with that situation — two candidates, so the
resolver refuses and records unresolved, because the source named two and
chose neither. In the 1909 transcription that is exactly what happens.&lt;/p&gt;
&lt;p&gt;The scan carries the same sentence with two words damaged. &lt;code&gt;purée&lt;/code&gt; comes
through as &lt;code&gt;pur^e&lt;/code&gt;, which changes nothing, because &lt;code&gt;tomato&lt;/code&gt; is the candidate
and it survives intact. And &lt;code&gt;velouté&lt;/code&gt; comes through as &lt;code&gt;velout^&lt;/code&gt;.&lt;/p&gt;
&lt;pre&gt;&lt;code class="language-console"&gt;$ uv run saucier show aurore-sauce --source escoffier-1907
AURORE SAUCE
entry 60, line 2503, ocr of escoffier-1907
  term  AURORE SAUCE  [en]  aurore-sauce
  parent  tomato
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;One candidate becomes unreadable. The other survives. The resolver now sees a
single unambiguous statement and records &lt;strong&gt;tomato&lt;/strong&gt; — confidently, and
wrongly.&lt;/p&gt;
&lt;p&gt;&lt;img src="https://alberto.codes/saucier-aurore-ambiguity.svg" alt="Two panels. On the left, escoffier-1909, the transcription, entry 60 line 2095: the sentence reads &amp;quot;boiling velouté ... very red tomato purée&amp;quot;, two candidates are stated, velouté and tomato, and the parent is unresolved because the source named two and chose neither. On the right, escoffier-1907, the ocr scan, entry 60 line 2503: the same sentence reads &amp;quot;boiling velout^ ... very red tomato pur^e&amp;quot;, only one candidate is stated because velouté is unreadable and tomato survives, and the parent is recorded as tomato — unambiguous, well formed, provenanced, and wrong. Both sides connect to a single note: the scan did not add noise to this record, it removed the ambiguity that was the reason for the honest answer, and nothing downstream can tell." /&gt;&lt;/p&gt;
&lt;p&gt;The scan did not add noise to that record. &lt;strong&gt;It removed the ambiguity that
was the reason for the honest answer&lt;/strong&gt;, and converted a correct refusal into a
confident wrong claim. And nothing downstream can tell: the record is well
formed, it validates, it carries a source id and an entry and a real line
number you can go and read. It is exactly the kind of output post #1 warned
about — plausible, checkable-looking, and false — arriving this time with no
model anywhere near it.&lt;/p&gt;
&lt;p&gt;Now notice what I just did to tell you that. I read &lt;code&gt;velout^&lt;/code&gt; as &lt;code&gt;velouté&lt;/code&gt;,
and I did it using French orthography and the other witness — evidence from
outside the document being read, which is exactly the evidence this project
forbids its own code to use. I cannot prove from the scan alone that
Escoffier printed &lt;code&gt;velouté&lt;/code&gt; there in 1907. I am confident he did, and my
confidence is a reader's judgement rather than a record's.&lt;/p&gt;
&lt;p&gt;That judgement is the adjudication this post keeps saying the tool cannot
perform. It turns out to exist. It exists in me, it is not written down
anywhere, and nothing in the catalogue is entitled to it.&lt;/p&gt;
&lt;p&gt;There is a one-character fix available and the project refuses to make it.
Repairing &lt;code&gt;velout^&lt;/code&gt; back to &lt;code&gt;velouté&lt;/code&gt; requires evidence from outside the
document being read: French orthography, or the other witness. The separator
repairs were allowed precisely because their evidence is internal — a line
reading &lt;code&gt;126-- MAYONNAISE SAUCE&lt;/code&gt; looks exactly like the thousands of
undamaged headings in the same file, and the repair changes no recorded byte, because
the separator never reaches a term. &lt;code&gt;QRIBICHE SAUCE&lt;/code&gt; is still recorded as
&lt;code&gt;QRIBICHE SAUCE&lt;/code&gt;. That is
&lt;a href="https://github.com/Alberto-Codes/saucier/blob/v0.3.0/docs/adr/0013-repair-structure-never-content.md"&gt;ADR-0013&lt;/a&gt;:
&lt;strong&gt;repair the punctuation that delimits a record, never the characters that
constitute one.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;The guard on that rule was measured rather than chosen, which I want to note
because it is the part people skip. Requiring the whole title in capitals was
tried first and held out two real headings, including the Chasseur one above.
Four opening capitals admits 57 lines in the scan, and every one of them is a
heading.&lt;/p&gt;
&lt;h2&gt;A second rule, and it is the transferable one&lt;/h2&gt;
&lt;p&gt;Until this week the series had exactly one epistemic rule: &lt;strong&gt;stated, never
inferred&lt;/strong&gt;. The catalogue records what the source says, an unresolved parent
stays unresolved rather than being guessed, and everything in the first two
posts descends from that.&lt;/p&gt;
&lt;p&gt;It now has a second one, and it is independent rather than a refinement:
&lt;strong&gt;observed, never assumed.&lt;/strong&gt;
&lt;a href="https://github.com/Alberto-Codes/saucier/blob/v0.3.0/docs/adr/0014-a-damaged-witness-cannot-establish-absence.md"&gt;ADR-0014&lt;/a&gt;
puts it plainly — absence is only observable through an instrument that can
see everything present, and where the instrument has a measured blind spot,
absence is unobservable and is not reported as observed.&lt;/p&gt;
&lt;p&gt;The two rules govern different questions. The first is about what you may
read from a source. The second is about what you may conclude from what you
read. A comparison involving a scanned witness now reports &lt;code&gt;unmatched&lt;/code&gt;, which
says the diff found no counterpart and says nothing about what either book
contains. Between two proofread texts, &lt;code&gt;added&lt;/code&gt; and &lt;code&gt;removed&lt;/code&gt; survive — the
rule is about damage, not about comparison.&lt;/p&gt;
&lt;p&gt;And the blind spot is printed beside the counts rather than in a footnote,
so no reader sees how many rows the diff found without also seeing how much
of the source it could not read. Today that is &lt;strong&gt;284 entries&lt;/strong&gt; of the 1907
scan that the parser still cannot see, against 2,679 it can.&lt;/p&gt;
&lt;p&gt;That number is the honest reason the next release exists, and it is the thing
that has to shrink before an absence claim comes back. Restoring &lt;code&gt;added&lt;/code&gt; for
a scanned witness is not a matter of a better comparison or a kinder
threshold. It requires reading the entries the parser currently cannot read,
and then measuring what is left.&lt;/p&gt;
&lt;h2&gt;What I am not claiming&lt;/h2&gt;
&lt;p&gt;Eighteen rows report a parent disagreement between the two witnesses. None of
them is a finding. Every one carries &lt;code&gt;ocr-suspected&lt;/code&gt;, which says a lost
candidate explains it as well as a revision does, and the diff adjudicates
none of them. Adjudicating one means putting a human eye on both lines — and
for the disputed spans, on the photographed page itself, because the question
is what the ink says rather than what the text file says. That capability
does not exist here yet.&lt;/p&gt;
&lt;p&gt;So there is no claim in this post about what Escoffier changed between 1907
and 1909. There is a catalogue of two witnesses, a diff that names what it
cannot distinguish, and a measured statement of how much of one witness
remains unread.&lt;/p&gt;
&lt;h2&gt;What breaks next&lt;/h2&gt;
&lt;p&gt;Post #2 promised the storage layer failing, and this is not that post either.
The reason is the same as the reason for everything above: a record that
named the wrong book had to stop naming the wrong book before anything got
built on top of it.&lt;/p&gt;
&lt;p&gt;The bill keeps growing while I do not pay it. There are two catalogues now,
&lt;code&gt;saucier diff&lt;/code&gt; loads both of them in full to compare them, and closing that
284-entry blind spot means reading a witness against its own page images —
which is a third artefact per entry, and the JSON file is still the entire
storage layer.&lt;/p&gt;
&lt;p&gt;Everything here reproduces from the tag:&lt;/p&gt;
&lt;pre&gt;&lt;code class="language-console"&gt;$ git clone https://github.com/Alberto-Codes/saucier
$ cd saucier &amp;amp;&amp;amp; git checkout v0.3.0
$ uv sync &amp;amp;&amp;amp; uv run saucier parse
$ uv run saucier diff escoffier-1907 escoffier-1909
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;That prints two witnesses of one book — 124 sauces and 115 — twenty-seven
rows where the catalogues disagree, a blind spot of 284 entries, and not one
adjudicated difference between two printings of &lt;em&gt;A Guide to Modern Cookery&lt;/em&gt;.
If you can read both lines of a disputed row and tell me which is the scanner
and which is Escoffier changing his mind,
&lt;a href="https://github.com/Alberto-Codes/saucier/issues/new?template=extraction.yml"&gt;the issue template&lt;/a&gt;
asks for the entry number and the source lines. It is the same thing I would
have needed to catch the twenty, and the eight.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;The release:&lt;/strong&gt; &lt;a href="https://github.com/Alberto-Codes/saucier/releases/tag/v0.3.0"&gt;saucier v0.3.0&lt;/a&gt;
— the second witness and the diff that reads it: the 1907 first printing beside
the 1909 revision, 9 unmatched rows, 18 parent-changed, 35 ocr-suspected and
not one of them adjudicated, and a 284-entry blind spot the tag admits to. MIT.&lt;/p&gt;
</content:encoded>
    </item>
    <item>
      <title>The 2-bit label was 4.5 bits inside. My 16 GiB card could tell.</title>
      <link>https://alberto.codes/blog/2026-08-22-the-2-bit-label-was-4-5-bits-inside</link>
      <guid isPermaLink="true">https://alberto.codes/blog/2026-08-22-the-2-bit-label-was-4-5-bits-inside</guid>
      <pubDate>Sat, 22 Aug 2026 00:00:00 GMT</pubDate>
      <description>The smallest 2-bit-labeled build of Nemotron 3.5 Lightning is 17.54 GiB — it doesn't fit a 16 GiB card, and its label names 12 of 417 tensors. I measured the model stack by stack instead, and got a 15.76 GiB pack that serves fully on-card at 16k context and beats the shelf's build on both damage metrics, while 1.78 GiB smaller.</description>
      <content:encoded>&lt;p&gt;I wanted &lt;a href="https://huggingface.co/nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16"&gt;NVIDIA's Nemotron 3.5 Lightning&lt;/a&gt; — a 30-billion-parameter mixture-of-experts model — running entirely on a 16 GiB card. The smallest GGUF on the shelf is labeled &lt;code&gt;IQ2_XXS&lt;/code&gt;, a 2-bit-class quantization, about as aggressively crushed as llama.cpp gets.&lt;/p&gt;
&lt;p&gt;It is &lt;strong&gt;17.54 GiB&lt;/strong&gt;. It does not fit.&lt;/p&gt;
&lt;p&gt;That's not sloppiness, and it's not anyone's dishonesty. The label is doing the only thing it can: naming &lt;strong&gt;12 of the file's 417 tensors&lt;/strong&gt;.&lt;/p&gt;
&lt;h2&gt;A label is not a recipe&lt;/h2&gt;
&lt;p&gt;&lt;a href="https://alberto.codes/blog/2026-08-11-i-couldnt-tell-my-quantized-model-from-the-baseline"&gt;Last time&lt;/a&gt;, the villain was the preset — a fixed rule that never measures the model in front of it. (That post also builds quantization from first principles; if &amp;quot;GGUF&amp;quot; and &amp;quot;2-bit&amp;quot; are new words, start there.) This time the preset doesn't even get to run. llama.cpp's compact quant types work on rows whose width 256 divides, and this model's routed-expert tensors are 2688 and 1856 columns wide. 256 divides neither. So the quantizer falls back, silently, and rewrites every one of those tensors to &lt;code&gt;IQ4_NL&lt;/code&gt; at &lt;strong&gt;4.5 bits per weight&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;This model stores the 128 routed experts of one projection in one layer as a single tensor — an &lt;em&gt;expert stack&lt;/em&gt;. Twenty-three MoE layers, an up and a down projection each: 46 stacks, and those &lt;strong&gt;46 stacks hold 93% of the parameters&lt;/strong&gt;. The fallback quietly rewrites all 46. Ask for the most aggressive 2-bit build and you get 4.5-bit experts wearing a 2-bit name tag.&lt;/p&gt;
&lt;p&gt;&lt;img src="https://alberto.codes/vramfit-label-vs-file.svg" alt="A bar of Nemotron 3.5 Lightning by parameter share. 93% of it is the 46 expert stacks, and the shelf's build silently rewrites every one to IQ4_NL at 4.5 bits per weight. The IQ2_XXS label names 12 of the build's 417 tensors, and not one expert stack is among them. The result is a 2-bit-class label on a 17.54 GiB file." /&gt;&lt;/p&gt;
&lt;p&gt;It's your abuelita's recipe again, one town over: the jar says what the jar has always said, and the kitchen has been quietly substituting an ingredient for years because the real one doesn't fit the pan. The dish still works. But if your table only seats 16 GiB, the label will not warn you — the smallest thing on the shelf simply doesn't fit, and the shelf's lower-labeled rungs are made of the same locked-out types, so they land at essentially the same size.&lt;/p&gt;
&lt;p&gt;The person who published that build, &lt;a href="https://huggingface.co/bartowski/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-GGUF"&gt;bartowski&lt;/a&gt;, did nothing wrong — the fallback belongs to the quantizer, and it runs without asking. This post exists because the way past it is the same move as last time: stop reading labels, start measuring.&lt;/p&gt;
&lt;h2&gt;Measure the stacks, not the layers&lt;/h2&gt;
&lt;p&gt;&lt;a href="https://github.com/Alberto-Codes/vramfit"&gt;vramfit&lt;/a&gt; built a sensitivity map keyed to what this model actually is. Not per-layer — per &lt;em&gt;expert stack&lt;/em&gt;, because a stack is the unit a GGUF pack assigns a precision to. Each of the 46 stacks was quantized alone, at each candidate precision, and the damage recorded: how far the model's output distribution shifts from the bf16 reference when only that group is crushed. 46 stacks × 2 candidate widths = 92 measured cells, each with its run-log line — wall clock, memory high-water mark, the number itself.&lt;/p&gt;
&lt;p&gt;Two full scans ran, and the difference between them matters later: one weighted its 4-bit fits with the same importance matrix the shelf's build used, and one ran unassisted. &lt;a href="https://huggingface.co/datasets/Alberto-Codes/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-sensitivity-maps"&gt;Both maps are published&lt;/a&gt;, run logs beside them.&lt;/p&gt;
&lt;p&gt;The dense weights — attention, Mamba-2, shared experts, embeddings, the output head — never entered the map at all. They're &lt;strong&gt;7% of the parameters&lt;/strong&gt;, which makes generosity nearly free: the recipe pins every quantizable dense class at 8-bit and spends the real decisions where 93% of the model lives.&lt;/p&gt;
&lt;h2&gt;The recipe a 16 GiB budget buys&lt;/h2&gt;
&lt;p&gt;The solver walks the map greedily by damage per byte saved, under the budget. What came out:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;11 expert stacks at 2.25 bits per weight&lt;/strong&gt; (&lt;code&gt;Q2_0&lt;/code&gt;) — all of them &lt;code&gt;down_proj&lt;/code&gt; stacks, the ones the map priced cheapest, landing spread across the depth from layer 22 on rather than clustered. The figure marks each one; the recipe records them exactly.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;The other 35 stacks at &lt;code&gt;Q4_0&lt;/code&gt;&lt;/strong&gt; — 4.5 bits, same real rate as the shelf's fallback.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;118 dense groups at &lt;code&gt;Q8_0&lt;/code&gt;&lt;/strong&gt;, and 46 groups passed through at F16 — the Mamba-2 convolutions and router gates, classes llama.cpp's quantizer never touches.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;img src="https://alberto.codes/vramfit-16gib-recipe.svg" alt="Two recipes for the same 46 expert stacks. The shelf's build spends 4.5 bits on all 46 via the fallback and carries an MTP block, landing at 17.54 GiB. This pack keeps 35 stacks at Q4_0 and demotes 11 down_proj stacks to Q2_0 at 2.25 bits, spread across the depth, landing at 15.76 GiB." /&gt;&lt;/p&gt;
&lt;p&gt;&lt;code&gt;Q2_0&lt;/code&gt; is the reason this recipe is possible at all: a plain block format — one scale per block, no super-block, no 256-wide requirement — that llama.cpp merged on 2026-07-07. It gives the solver a real 2.25-bit rung on tensors the compact types refuse. (It also means the file needs a llama.cpp build that carries that merge — a build older than it refuses the file. The serve test ran b10326.)&lt;/p&gt;
&lt;p&gt;The recipe isn't a magic constant. It records the full solve — the budget bytes, the nine pins, the 11-step demotion trace — so the solve replays from the artifact, and the pack rebuilds this file's bytes on one machine from the base checkpoint. Across machines the byte count can differ by tens of bytes (the GGUF metadata stores the imatrix path); the recipe, not the checksum, is the identity.&lt;/p&gt;
&lt;h2&gt;fit16gib is a contract, not a boast&lt;/h2&gt;
&lt;p&gt;The name on the repo says &lt;code&gt;fit16gib&lt;/code&gt;, and I want to be precise about what that claims, because &amp;quot;it fits&amp;quot; is exactly the kind of unweighed statement this project exists to replace.&lt;/p&gt;
&lt;p&gt;&lt;img src="https://alberto.codes/vramfit-16gib-budget.svg" alt="A 16 GiB card splits into a 15.776 GiB weight budget and a 0.224 GiB runtime reserve for KV cache, recurrent state, and compute. The packed file lands at 15.760 GiB, 16.09 MiB under the budget. The margin holds to 8 parallel sequences." /&gt;&lt;/p&gt;
&lt;p&gt;The claim is: this file loads &lt;strong&gt;fully offloaded&lt;/strong&gt; on a 16 GiB card, holds &lt;strong&gt;16k of context&lt;/strong&gt;, and generates. It's a measured serve result under a stated configuration, not a promise about every runtime.&lt;/p&gt;
&lt;p&gt;The test: a hard ballast cap held an RTX 4090 to 16,383 MiB visible, and llama.cpp b10326 (Vulkan) loaded the file with every layer offloaded — 53 of 53, 15,774.00 MiB of weights on the device, with the 357.00 MiB token embedding host-mapped, as llama.cpp always keeps it. &lt;code&gt;llama-server&lt;/code&gt; answered a completion request from inside that envelope: 16,157.88 MiB of device buffers against the 16,383 visible. The buffers total more than the 0.224 GiB reserve because the server ran four slots and recurrent state grows per sequence — which is exactly the growth the margin exists to absorb. The published build cannot take this test: its 17.54 GiB of weights exceed the card before the first buffer allocates.&lt;/p&gt;
&lt;p&gt;And if that 0.224 GiB reserve looks suspiciously small beside the multi-GiB context reserve of &lt;a href="https://alberto.codes/blog/2026-08-11-i-couldnt-tell-my-quantized-model-from-the-baseline"&gt;publication #1&lt;/a&gt;: this model is a Mamba-2 hybrid with six attention layers, so 16k of KV cache is 96.00 MiB. The reserve is measured buffers rounded up, not a guessed overhead.&lt;/p&gt;
&lt;p&gt;The contract's edges, stated plainly:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;The margin absorbs recurrent-state growth to &lt;strong&gt;8 parallel sequences&lt;/strong&gt;. Above 8, the claim is off.&lt;/li&gt;
&lt;li&gt;The test's 16,383 MiB is a ballast-capped 4090, not your card. A 16 GiB card driving a display, a different backend, or a runtime with different buffer sizes moves the envelope — the claim is the stated configuration, and the card states it.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;No tokens-per-second figure appears here or on the card.&lt;/strong&gt; The serve ran on a VRAM-capped 4090, and a decode number from that method would read 1.4 to 3.5 times higher than real 16 GiB silicon delivers. Publishing it would be publishing a flattering lie.&lt;/li&gt;
&lt;li&gt;A 16 GiB owner can already run &lt;em&gt;bigger&lt;/em&gt; builds today by spilling part of the weights to CPU and accepting slower decode. That is a real alternative, and the project has not measured its speed cost. This pack is the option that keeps every weight on the card; whether the trade favors it on your workload is genuinely unsettled.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2&gt;The scoreboard&lt;/h2&gt;
&lt;p&gt;Both damage metrics were ruled before the measurement ran, read together from one instrument: the &lt;strong&gt;perplexity ratio&lt;/strong&gt; (how much worse than the f16 original the pack predicts held-out text) and &lt;strong&gt;mean KL divergence&lt;/strong&gt; (how far its output probabilities drift from the original's). Held-out WikiText-2, 594 chunks — my 15.76 GiB pack against the shelf's 17.54 GiB build:&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;PPL / f16 ↓&lt;/th&gt;
&lt;th&gt;Mean KLD ↓&lt;/th&gt;
&lt;th&gt;Same top ↑&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;This pack&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;1.1611&lt;/td&gt;
&lt;td&gt;0.2043&lt;/td&gt;
&lt;td&gt;83.13%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;IQ2_XXS&lt;/code&gt; (the shelf)&lt;/td&gt;
&lt;td&gt;1.3209&lt;/td&gt;
&lt;td&gt;0.3703&lt;/td&gt;
&lt;td&gt;76.09%&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;Both ruled metrics at once: a perplexity ratio 0.16 lower and &lt;strong&gt;44.8% lower mean KL divergence&lt;/strong&gt;, at &lt;strong&gt;1.78 GiB fewer bytes&lt;/strong&gt;. Top-token agreement rides along uninvited — it wasn't ruled, and in publication #1 it was the metric that beat &lt;em&gt;me&lt;/em&gt; — so it reports here, but it doesn't rank. Last time the win was 2.9% on one metric at matched size. This one isn't close.&lt;/p&gt;
&lt;p&gt;Then the task slice — five benchmarks fixed before any run, full evaluation splits, both models on the same lane. A delta inside the combined standard error reports as a tie; Winogrande clears that bar by a whisker:&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Task&lt;/th&gt;
&lt;th&gt;This pack&lt;/th&gt;
&lt;th&gt;The shelf's build&lt;/th&gt;
&lt;th&gt;Verdict&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;MMLU (5-shot)&lt;/td&gt;
&lt;td&gt;0.7651&lt;/td&gt;
&lt;td&gt;0.6848&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;ahead, +8.0 points at 16.1σ&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;HellaSwag (10-shot)&lt;/td&gt;
&lt;td&gt;0.8038&lt;/td&gt;
&lt;td&gt;0.7652&lt;/td&gt;
&lt;td&gt;ahead (6.7σ)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GSM8K (5-shot)&lt;/td&gt;
&lt;td&gt;0.7839&lt;/td&gt;
&lt;td&gt;0.7627&lt;/td&gt;
&lt;td&gt;ahead (1.3σ)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Winogrande (5-shot)&lt;/td&gt;
&lt;td&gt;0.7443&lt;/td&gt;
&lt;td&gt;0.7261&lt;/td&gt;
&lt;td&gt;ahead (1.04σ, barely)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;ARC-Challenge (25-shot)&lt;/td&gt;
&lt;td&gt;0.6630&lt;/td&gt;
&lt;td&gt;0.6715&lt;/td&gt;
&lt;td&gt;tie (0.4σ)&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;Four leads and a tie, from the smaller file. The one nominal deficit prints with its error bar.&lt;/p&gt;
&lt;p&gt;One cell I won't fill: the shelf build's bits-per-parameter. Mine is 4.287; its bytes include a block mine omits (next section), so the division would run over different weights. An empty cell is more honest than a wrong one.&lt;/p&gt;
&lt;h2&gt;The handicap ran the wrong way&lt;/h2&gt;
&lt;p&gt;The cleanest objection: &lt;em&gt;you used a different importance matrix.&lt;/em&gt; No — both packs consumed &lt;strong&gt;the same one&lt;/strong&gt;, bartowski's, 185 entries over 822 chunks. It stays &lt;a href="https://huggingface.co/bartowski/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-GGUF"&gt;in his repository&lt;/a&gt;, linked from my card at a pinned revision with its hash, because it's his work and a link with credit is the right way to use it.&lt;/p&gt;
&lt;p&gt;And the assistance ran &lt;em&gt;against&lt;/em&gt; me. His build quantizes &lt;strong&gt;91.53%&lt;/strong&gt; of its bytes with the matrix's help. Mine: &lt;strong&gt;74.44%&lt;/strong&gt; — no type takes an assisted fit at 2.25 bits on these row widths, and the 8-bit quantizer discards the matrix outright. The pack wins carrying less of the shared advantage.&lt;/p&gt;
&lt;p&gt;Which raises the attribution question the second scan exists to answer: is the win the map's, or the matrix's? The unassisted map — measured without the importance matrix — agrees with the assisted one through every rank the solve reads, and yields the identical placement. The win credits the damage ranking, not the borrowed weighting.&lt;/p&gt;
&lt;h2&gt;Placement was a decision, not a default&lt;/h2&gt;
&lt;p&gt;One more claim earned its keep before publication: that &lt;em&gt;where&lt;/em&gt; the 11 cheap stacks land matters, and the map knows where.&lt;/p&gt;
&lt;p&gt;The same campaign packed and measured &lt;strong&gt;nine placements of the identical width mix&lt;/strong&gt; — the map-ranked one, a spread-map probe, three blind draws, a spread-matched control, a measured-map arm, a class-wise arm, and one deliberately inverted to sit on the stacks the map prices dearest. Same 11-at-2.25, 35-at-4.5 spend, nine different answers to &amp;quot;which eleven.&amp;quot; The map-ranked placement measured the least damage of the nine; the worst arm measured &lt;strong&gt;2.7 times&lt;/strong&gt; as much.&lt;/p&gt;
&lt;p&gt;Same bytes, same widths, 2.7× spread in damage. Allocation decides — and the map ranked it correctly before any of the nine files existed.&lt;/p&gt;
&lt;h2&gt;What I'm still not claiming&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;This pack is missing a block the shelf's build has.&lt;/strong&gt; The base checkpoint ships multi-token-prediction layers; my conversion dropped them (&lt;code&gt;--no-mtp&lt;/code&gt;), so speculative decoding off that block is unavailable from this file. The comparator carries its MTP block at Q4_0 — which is also part of why it's bigger. If MTP-based speculation matters to your serving stack, that's a real difference, not a detail.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;No speed claim survives this post.&lt;/strong&gt; Not mine — the capped-4090 method inflates it. Not the CPU-offload alternative's — nobody measured it. The honest sentence is: this is the only &lt;em&gt;fully-on-card&lt;/em&gt; option I can prove exists at 16 GiB and 16k context, and how much that's worth over offloading is an open question I've chosen to leave open in public rather than settle with a number I don't have.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Don't take the damage numbers on tour.&lt;/strong&gt; The sensitivity values are one scan's measurements in one frame. They rank stacks within this map, and that's all they do — comparing them across scans, calibration sets, or models is meaningless, and the dataset card says so in bold. Rank packed models by measured quality at a fixed model and budget, never by raw damage.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;And the ledger keeps growing.&lt;/strong&gt; Every claim above traces to &lt;a href="https://github.com/Alberto-Codes/vramfit/blob/main/docs/explanation/evaluating-packed-models.md"&gt;the evaluation record&lt;/a&gt;, the same public trail the &lt;a href="https://alberto.codes/blog/2026-08-11-i-couldnt-tell-my-quantized-model-from-the-baseline"&gt;first publication&lt;/a&gt; started — including the arms that lost. The going standard for a published quant is still a checksum and a file size. A checksum proves the file is the file. These documents are the argument that it's any good.&lt;/p&gt;
&lt;h2&gt;Where it lives&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;The model:&lt;/strong&gt; &lt;a href="https://huggingface.co/Alberto-Codes/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-fit16gib-GGUF"&gt;NVIDIA-Nemotron-3.5-Lightning-30B-A3B-fit16gib-GGUF&lt;/a&gt; — a quantization of &lt;a href="https://huggingface.co/nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16"&gt;NVIDIA's Nemotron 3.5 Lightning 30B-A3B&lt;/a&gt;, under OpenMDW 1.1. The card carries the recipe, the serve log, the eval sidecars, and the reproduce commands.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;The measurements:&lt;/strong&gt; &lt;a href="https://huggingface.co/datasets/Alberto-Codes/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-sensitivity-maps"&gt;the sensitivity-map dataset&lt;/a&gt; — both scans, run logs beside them. Solve your own card's budget against it.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;The how-to:&lt;/strong&gt; &lt;a href="https://alberto.codes/blog/2026-08-15-fit-a-model-to-the-gpu-you-actually-have"&gt;fit a model to the GPU you actually have&lt;/a&gt; — the command-by-command version, for your own card and model.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;The tool:&lt;/strong&gt; &lt;a href="https://github.com/Alberto-Codes/vramfit/releases/tag/v0.3.0"&gt;vramfit v0.3.0&lt;/a&gt; — the tag that solved and packed this build: the stack-keyed scan, the spread placement rule, and the q0/q0-imx meters. MIT.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;A week ago I &lt;a href="https://alberto.codes/blog/2026-08-15-a-different-ceiling-is-a-different-recipe"&gt;argued&lt;/a&gt; that a different ceiling is a different recipe, on paper — and that for the 49B at 16 GiB, the honest answer was no dish at all. That model still doesn't fit a 16 GiB card. This is a different model, solved for that smaller ceiling from the start: a locked-out quant family, a stack-keyed map, and a recipe that looks nothing like last time's, because the model is nothing like last time's. The method didn't change. Measure, then solve.&lt;/p&gt;
</content:encoded>
    </item>
    <item>
      <title>Marrow Sauce finally has a parent. The parser gave up three answers to earn it.</title>
      <link>https://alberto.codes/blog/2026-08-21-marrow-sauce-finally-has-a-parent</link>
      <guid isPermaLink="true">https://alberto.codes/blog/2026-08-21-marrow-sauce-finally-has-a-parent</guid>
      <pubDate>Fri, 21 Aug 2026 00:00:00 GMT</pubDate>
      <description>Post two of the saucier series lets a parent be any catalogued preparation, not only one of the five mothers. Derived rises from 29 to 50 and every new derivation quotes a name the source wrote. The same rule dissolves three derivations the old parser recorded, one of them caught by peer review — and Bordelaise stays unresolved on purpose, because a term that encodes a derivation is not a statement of one.</description>
      <content:encoded>&lt;p&gt;Post #1 ended with a promise: the next post would be the JSON file failing,
because that is the only way I am willing to introduce a database. This is not
that post. The chain resolver was promised in the same breath, it costs no
infrastructure, and it cut in line. The storage failure is still coming — and
by the end of this post it will be closer than it was, for reasons the
resolver itself just demonstrated.&lt;/p&gt;
&lt;p&gt;Here is what changed, in one sentence: a preparation's parent may now be any
catalogued preparation, not only one of the five mothers. Here is what that
did to the census, at the tag that carries it, &lt;code&gt;v0.2.0&lt;/code&gt;:&lt;/p&gt;
&lt;pre&gt;&lt;code class="language-console"&gt;$ uv run saucier parse
source      escoffier-1907
mothers     bechamel, espagnole, hollandaise, tomato, veloute
sauces      124
derived     50 linked to a stated parent
unresolved  74 state no base in their prose
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;&lt;em&gt;Correction, 2026-09-01: that &lt;code&gt;escoffier-1907&lt;/code&gt; is the identifier this release
printed, and it was wrong. The file is the New and Revised Edition of January
1909, and its own title page says so. Everything else here is unaffected and
still reproduces at &lt;code&gt;v0.2.0&lt;/code&gt;.
&lt;a href="https://alberto.codes/blog/2026-09-01-i-added-a-second-copy-of-the-same-book"&gt;Post three&lt;/a&gt; is the
correction.&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;Twenty-nine became fifty. Still no model, no API key, no network — a regular
expression got better at reading, nothing more. And the number I want to show
you first is not the 24 derivations it added. It is the three it took away.&lt;/p&gt;
&lt;h2&gt;The promise, kept literally&lt;/h2&gt;
&lt;p&gt;Post #1 left Marrow Sauce as the standing embarrassment. Entry 45 opens:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Follow the proportions as indicated under &amp;quot;Sauce Bordelaise&amp;quot; (No. 32) for the
necessary quantity of this sauce, the Marrow Sauce being only a variety of
the Bordelaise.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;An author stating a derivation in plain English, and a record answering
&lt;code&gt;parent: null&lt;/code&gt; — because Bordelaise is not a mother, and the old rule only
resolved mothers. The fix was named in that post: resolve to any catalogued
sauce and walk the chain. Deterministic, checkable line by line, no model
required.&lt;/p&gt;
&lt;p&gt;That rule now exists, and:&lt;/p&gt;
&lt;pre&gt;&lt;code class="language-console"&gt;$ uv run saucier tree bordelaise
SAUCE BORDELAISE  [bordelaise]
└── MARROW SAUCE  (en)
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Readers of post #1 have seen this picture before. One arrow in it is different
now.&lt;/p&gt;
&lt;p&gt;&lt;img src="https://alberto.codes/saucier-unstated-chain-resolved.svg" alt="A chain of four sauces descending. At the top, Espagnole, entry 22 line 1392, the only one the source explicitly calls a mother. A dashed arrow marked &amp;quot;assumed, not stated&amp;quot; runs down to half-glaze, a reduction of Espagnole the book never spells out. Another dashed &amp;quot;assumed, not stated&amp;quot; arrow runs to Sauce Bordelaise, entry 32 line 1680, whose opening says &amp;quot;half-glaze&amp;quot; and never &amp;quot;Espagnole&amp;quot;, so its parent is still null. The last arrow, marked &amp;quot;stated outright, now resolved&amp;quot;, is solid: Marrow Sauce, entry 45 line 1895, which calls itself &amp;quot;only a variety of the Bordelaise&amp;quot;, now records parent: sauce-bordelaise." /&gt;&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Correction, 2026-09-03: this figure and the paragraph beneath it say the book never spells out that half glaze is a reduction of Espagnole. It does, at entry 23, line 1437, in the first sentence of an entry the parser at this tag could not see. The figure is left as printed, because it shows what the parser recorded. &lt;a href="https://alberto.codes/blog/2026-09-03-the-book-spells-it-out-at-line-1437"&gt;The post about line 1437&lt;/a&gt; is the correction.&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;Notice what did not change. Bordelaise's own parent is still null. Its opening
names shallots, red wine, mignonette pepper, thyme, bay, and half a pint of
half-glaze — and a &lt;em&gt;demi-glace&lt;/em&gt; is an Espagnole reduction, which every reader
of Escoffier knew and the book therefore never says. The resolver could have
special-cased it in one line. It did not, because a term that encodes a
derivation is not a statement of one, and the moment this parser records one
inference it stops being the baseline the eventual model gets measured
against. Bordelaise staying unresolved is not a gap in this release. It is the
bar the model post will have to clear, set in print for the second time.&lt;/p&gt;
&lt;h2&gt;Widening the net without catching ice cream&lt;/h2&gt;
&lt;p&gt;The obvious risk here has a precedent. Post #1's original census put vanilla
ice cream in the sauce catalogue because a substring test matched &lt;code&gt;hollandaise&lt;/code&gt;
inside &lt;code&gt;BOMBE HOLLANDAISE&lt;/code&gt;. The old parent rule had five candidate names to
match. This one has every name in the catalogue — well over a hundred, in two
languages, many of them containing each other. Widening the candidate set
widens the ways a match can be wrong. The whole job of this release was to
take the wider set without the wider wrongness.&lt;/p&gt;
&lt;p&gt;Three rules do that work, and each one exists because the corpus punished its
absence (&lt;a href="https://github.com/Alberto-Codes/saucier/blob/v0.2.0/docs/adr/0008-a-parent-may-be-any-catalogued-preparation.md"&gt;ADR-0008&lt;/a&gt;):&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;A statement is a whole run of words inside one sentence.&lt;/strong&gt; Folding an
opening paragraph flattens punctuation, and a matcher that joins words across
a full stop is the ice cream defect wearing a new costume. &lt;code&gt;tomatoes&lt;/code&gt; is
still not &lt;code&gt;tomato&lt;/code&gt;, and a name split across two sentences is not a name.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;A run inside the entry's own name states an ingredient, not a parent.&lt;/strong&gt;
Entry 138, &lt;code&gt;HORSE-RADISH SAUCE&lt;/code&gt; (line 3091), opens with &amp;quot;finely-rasped
horse-radish&amp;quot; — words that match the catalogued &lt;code&gt;HORSE-RADISH OR ALBERT SAUCE&lt;/code&gt;
(entry 119, line 2804). It is naming its own subject. It stays unresolved.
Mothers are exempt from this rule, because a mother is never an entry's own
subject — which is why &lt;code&gt;LENTEN ESPAGNOLE&lt;/code&gt; still resolves to Espagnole.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;A longer stated name shadows the shorter one inside it.&lt;/strong&gt; Entry 38,
Genevoise (line 1769), opens with &amp;quot;add one pint of Lenten Espagnole&amp;quot;. The word
&amp;quot;Espagnole&amp;quot; sits inside that name. Reading both would turn one clear statement
into a fake ambiguity; reading only the longer one records what the author
wrote. Under the old rule Genevoise resolved straight to the mother Espagnole —
which sounds right and is wrong. The source says Lenten Espagnole, entry 24,
line 1449, a catalogued preparation with its own stated parent. So the record
now holds a chain, every derivation in it stated:&lt;/p&gt;
&lt;pre&gt;&lt;code class="language-console"&gt;$ uv run saucier tree espagnole
BROWN SAUCE OR ESPAGNOLE  [espagnole]
├── LENTEN ESPAGNOLE  (fr)
│   └── GENEVOISE SAUCE  (en)
├── ORDINARY POIVRADE SAUCE  (en)
│   └── REFORM SAUCE  (en)
└── POIVRADE SAUCE FOR VENISON  (en)
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Escoffier's structure was never a five-way star. The tree finally has the
depth the book always had.&lt;/p&gt;
&lt;h2&gt;The three answers it gave back&lt;/h2&gt;
&lt;p&gt;Ambiguity still resolves to nothing: exactly one candidate in the opening
paragraph, or no parent. Post #1 made that rule sound like modesty. This
release shows what it is actually for — because with more candidates in play,
the rule started dissolving derivations the old parser had recorded.&lt;/p&gt;
&lt;p&gt;&lt;code&gt;ANDALOUSE SAUCE&lt;/code&gt; (entry 122, line 2855) was on the books as derived from the
mother Tomato. Its opening reads:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Take the required quantity of Mayonnaise sauce (No. 126) and add to it the
quarter of its volume of very red and concentrated tomato purée…&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;The old rule could only see mothers, found the word &amp;quot;tomato&amp;quot;, and recorded a
derivation — read off an ingredient purée. The new rule sees two candidate
names, Mayonnaise and tomato, and two candidates means no parent. The wider
net did not just add derivations. It exposed one of the old ones as exactly
the grilled-tomatoes defect post #1 was about, one rung further up.
Chaud-froid au vert-pré (entry 75, line 2258) went the same way: its opening
names the velouté &lt;em&gt;and&lt;/em&gt; the white Chaud-Froid sauce, the source chose
neither, so neither does the record.&lt;/p&gt;
&lt;p&gt;The third one took a human. Post #1's census bugs were caught by a review
tool; this release's was caught by a person reviewing the pull request, and
it is the same species of defect one layer deeper. The resolver bound each
mother to a catalogued entry by preferring the least qualified name — which
bound the mother velouté to &lt;code&gt;THICKENED VELOUTÉ&lt;/code&gt;, an alias of Allemande
(entry 27), instead of &lt;code&gt;ORDINARY VELOUTÉ SAUCE&lt;/code&gt; (entry 25). Ordinary
Chaud-froid (entry 73, line 2242) opens with &amp;quot;substituting Allemande Sauce
for the velouté&amp;quot; — two stated candidates, which the bad binding collapsed
into one, so the record answered velouté where the ambiguity rule owed it
nothing. The fix binds a mother to the first preparation in source order
that answers to its name, because the source states a base before its
derivatives — and entry 73 dissolved into the unresolved column, moving the
census from 51 to 50 before either number ever got published.&lt;/p&gt;
&lt;p&gt;A resolver that can only add derivations is a ratchet. This one handed back
three confident answers because the evidence for them stopped clearing the
bar, and that — more than the 24 additions — is why the additions are
believable.&lt;/p&gt;
&lt;p&gt;&lt;code&gt;SHRIMP SAUCE&lt;/code&gt; (entry 80, line 2322) still names fish velouté &amp;quot;or, failing
this, Béchamel&amp;quot;, still gets no parent, and is still the correct answer.&lt;/p&gt;
&lt;h2&gt;Every changed record, with receipts&lt;/h2&gt;
&lt;p&gt;Post #1 counted twenty-seven unresolved preparations that name another
catalogued sauce in their opening. The resolver records twenty-four new
derivations — the difference is the gap between naming and stating, which is
precisely the gap the three rules above enforce. Add the three dissolved
derivations and Genevoise's corrected parent and 28 records changed, every
one checkable with &lt;code&gt;sed -n&lt;/code&gt; against the committed corpus:&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;entry&lt;/th&gt;
&lt;th&gt;line&lt;/th&gt;
&lt;th&gt;preparation&lt;/th&gt;
&lt;th&gt;before&lt;/th&gt;
&lt;th&gt;after&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;38&lt;/td&gt;
&lt;td&gt;1769&lt;/td&gt;
&lt;td&gt;GENEVOISE SAUCE&lt;/td&gt;
&lt;td&gt;espagnole&lt;/td&gt;
&lt;td&gt;lenten-espagnole&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;45&lt;/td&gt;
&lt;td&gt;1895&lt;/td&gt;
&lt;td&gt;MARROW SAUCE&lt;/td&gt;
&lt;td&gt;unresolved&lt;/td&gt;
&lt;td&gt;sauce-bordelaise&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;47&lt;/td&gt;
&lt;td&gt;1920&lt;/td&gt;
&lt;td&gt;PÉRIGUEUX SAUCE&lt;/td&gt;
&lt;td&gt;unresolved&lt;/td&gt;
&lt;td&gt;madeira-sauce&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;59&lt;/td&gt;
&lt;td&gt;2087&lt;/td&gt;
&lt;td&gt;ANCHOVY SAUCE&lt;/td&gt;
&lt;td&gt;unresolved&lt;/td&gt;
&lt;td&gt;normande-sauce&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;68&lt;/td&gt;
&lt;td&gt;2185&lt;/td&gt;
&lt;td&gt;CAPER SAUCE&lt;/td&gt;
&lt;td&gt;unresolved&lt;/td&gt;
&lt;td&gt;butter-sauce&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;70&lt;/td&gt;
&lt;td&gt;2202&lt;/td&gt;
&lt;td&gt;MUSHROOM SAUCE&lt;/td&gt;
&lt;td&gt;unresolved&lt;/td&gt;
&lt;td&gt;allemande-sauce&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;73&lt;/td&gt;
&lt;td&gt;2242&lt;/td&gt;
&lt;td&gt;ORDINARY CHAUD-FROID SAUCE&lt;/td&gt;
&lt;td&gt;veloute&lt;/td&gt;
&lt;td&gt;unresolved&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;75&lt;/td&gt;
&lt;td&gt;2258&lt;/td&gt;
&lt;td&gt;CHAUD-FROID SAUCE, AU VERT-PRÉ&lt;/td&gt;
&lt;td&gt;veloute&lt;/td&gt;
&lt;td&gt;unresolved&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;82&lt;/td&gt;
&lt;td&gt;2347&lt;/td&gt;
&lt;td&gt;DIPLOMATE SAUCE&lt;/td&gt;
&lt;td&gt;unresolved&lt;/td&gt;
&lt;td&gt;normande-sauce&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;83&lt;/td&gt;
&lt;td&gt;2354&lt;/td&gt;
&lt;td&gt;HERB SAUCE&lt;/td&gt;
&lt;td&gt;unresolved&lt;/td&gt;
&lt;td&gt;white-wine-sauce&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;84&lt;/td&gt;
&lt;td&gt;2362&lt;/td&gt;
&lt;td&gt;GOOSEBERRY SAUCE&lt;/td&gt;
&lt;td&gt;unresolved&lt;/td&gt;
&lt;td&gt;butter-sauce&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;86&lt;/td&gt;
&lt;td&gt;2390&lt;/td&gt;
&lt;td&gt;OYSTER SAUCE&lt;/td&gt;
&lt;td&gt;unresolved&lt;/td&gt;
&lt;td&gt;normande-sauce&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;88&lt;/td&gt;
&lt;td&gt;2406&lt;/td&gt;
&lt;td&gt;JOINVILLE SAUCE&lt;/td&gt;
&lt;td&gt;unresolved&lt;/td&gt;
&lt;td&gt;normande-sauce&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;90&lt;/td&gt;
&lt;td&gt;2428&lt;/td&gt;
&lt;td&gt;MARINIÈRE SAUCE&lt;/td&gt;
&lt;td&gt;unresolved&lt;/td&gt;
&lt;td&gt;bercy-sauce&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;94&lt;/td&gt;
&lt;td&gt;2468&lt;/td&gt;
&lt;td&gt;MUSTARD SAUCE&lt;/td&gt;
&lt;td&gt;unresolved&lt;/td&gt;
&lt;td&gt;butter-sauce&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;100&lt;/td&gt;
&lt;td&gt;2560&lt;/td&gt;
&lt;td&gt;ORIENTAL SAUCE&lt;/td&gt;
&lt;td&gt;unresolved&lt;/td&gt;
&lt;td&gt;american-sauce&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;103&lt;/td&gt;
&lt;td&gt;2587&lt;/td&gt;
&lt;td&gt;REGENCY SAUCE&lt;/td&gt;
&lt;td&gt;unresolved&lt;/td&gt;
&lt;td&gt;allemande-sauce&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;107&lt;/td&gt;
&lt;td&gt;2661&lt;/td&gt;
&lt;td&gt;VENETIAN SAUCE&lt;/td&gt;
&lt;td&gt;unresolved&lt;/td&gt;
&lt;td&gt;white-wine-sauce&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;108&lt;/td&gt;
&lt;td&gt;2672&lt;/td&gt;
&lt;td&gt;VILLEROY SAUCE&lt;/td&gt;
&lt;td&gt;unresolved&lt;/td&gt;
&lt;td&gt;allemande-sauce&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;109&lt;/td&gt;
&lt;td&gt;2681&lt;/td&gt;
&lt;td&gt;VILLEROY SOUBISEE SAUCE&lt;/td&gt;
&lt;td&gt;unresolved&lt;/td&gt;
&lt;td&gt;allemande-sauce&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;114&lt;/td&gt;
&lt;td&gt;2749&lt;/td&gt;
&lt;td&gt;CELERY SAUCE&lt;/td&gt;
&lt;td&gt;unresolved&lt;/td&gt;
&lt;td&gt;cream-sauce&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;116&lt;/td&gt;
&lt;td&gt;2777&lt;/td&gt;
&lt;td&gt;FENNEL SAUCE&lt;/td&gt;
&lt;td&gt;unresolved&lt;/td&gt;
&lt;td&gt;butter-sauce&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;119&lt;/td&gt;
&lt;td&gt;2804&lt;/td&gt;
&lt;td&gt;HORSE-RADISH OR ALBERT SAUCE&lt;/td&gt;
&lt;td&gt;unresolved&lt;/td&gt;
&lt;td&gt;butter-sauce&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;120&lt;/td&gt;
&lt;td&gt;2823&lt;/td&gt;
&lt;td&gt;REFORM SAUCE&lt;/td&gt;
&lt;td&gt;unresolved&lt;/td&gt;
&lt;td&gt;ordinary-poivrade-sauce&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;122&lt;/td&gt;
&lt;td&gt;2855&lt;/td&gt;
&lt;td&gt;ANDALOUSE SAUCE&lt;/td&gt;
&lt;td&gt;tomato&lt;/td&gt;
&lt;td&gt;unresolved&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;137&lt;/td&gt;
&lt;td&gt;3083&lt;/td&gt;
&lt;td&gt;OXFORD SAUCE&lt;/td&gt;
&lt;td&gt;unresolved&lt;/td&gt;
&lt;td&gt;cumberland-sauce&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2412&lt;/td&gt;
&lt;td&gt;38683&lt;/td&gt;
&lt;td&gt;SAUCE ORANGE&lt;/td&gt;
&lt;td&gt;unresolved&lt;/td&gt;
&lt;td&gt;apricot-sauce&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2414&lt;/td&gt;
&lt;td&gt;38696&lt;/td&gt;
&lt;td&gt;GREENGAGE OR MIRABELLE SAUCE&lt;/td&gt;
&lt;td&gt;unresolved&lt;/td&gt;
&lt;td&gt;apricot-sauce&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;Read down the &lt;em&gt;after&lt;/em&gt; column and the book's real middle layer appears: plain
Butter Sauce quietly picks up five children — caper, gooseberry, mustard,
fennel, and the Albert sauce. Normande gains four. And down in the dessert
chapter, Sauce Orange and the Greengage resolve to Apricot Sauce — which is
itself unresolved, so the tree now roots in a preparation that states no base
of its own:&lt;/p&gt;
&lt;pre&gt;&lt;code class="language-console"&gt;$ uv run saucier tree apricot-sauce
APRICOT SAUCE  [apricot-sauce]
├── SAUCE ORANGE  (fr)
└── GREENGAGE OR MIRABELLE SAUCE  (en)
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;That render is not an edge case someone forgot to reject. It is the data
model saying what the source says: these two derive from that one, and that
one keeps its counsel. Unresolved is a fact about the text, not a hole in the
tree.&lt;/p&gt;
&lt;p&gt;Two smaller guarantees ride along. Every new derivation carries the same
provenance as every old claim — source id, entry, line — which is what made
the table above checkable in an afternoon. And a preparation can never become
its own ancestor: a cycle is cleared entirely rather than broken by choosing
a derivation to keep, because choosing would be an arbitrary decision wearing
the costume of determinism, and this project already shipped one of those
(&lt;code&gt;SHRIMP SAUCE&lt;/code&gt;, resolved to Béchamel because B sorts before V, corrected in
post #1).&lt;/p&gt;
&lt;h2&gt;What this prices&lt;/h2&gt;
&lt;p&gt;The case for a model was priced in post #1 at 95 unresolved preparations. It
is now priced at 74, and the discount came from the cheapest possible vendor:
a stricter reading of what the source already says.&lt;/p&gt;
&lt;p&gt;That matters for how the eventual model gets judged. Every derivation a
deterministic pass can recover is one the model no longer gets credit for
recovering. What remains in the 74 is the genuinely hard residue — Bordelaise,
where the derivation hides inside the term &lt;em&gt;half-glaze&lt;/em&gt;; Shrimp, where the
author offered two bases and committed to neither; and the long tail of
preparations that simply never say. When a model finally reads those, its
score will mean something, because everything a regular expression could
honestly claim has already been claimed.&lt;/p&gt;
&lt;h2&gt;What it cost to write 28 fields&lt;/h2&gt;
&lt;p&gt;One more number, and it is the ending post #1 actually promised. Changing 28
parents meant rewriting the entire catalogue — all 124 records, serialized
back out as one JSON file, because that is the entire storage layer. This
release felt the cost that post #1 only predicted: the write is all-or-nothing,
the diff is the whole file, and every future rule that touches a handful of
records will pay the same full-file toll. The next post is that seam finally
tearing, and what replaces it — which is still the only way I am willing to
introduce a database.&lt;/p&gt;
&lt;p&gt;Everything here reproduces from the tag:&lt;/p&gt;
&lt;pre&gt;&lt;code class="language-console"&gt;$ git clone https://github.com/Alberto-Codes/saucier
$ cd saucier &amp;amp;&amp;amp; git checkout v0.2.0
$ uv sync &amp;amp;&amp;amp; uv run saucier parse
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;124 sauces. 50 derived, every derivation stated. 74 that keep their counsel,
on purpose. If you find a stated parent the resolver missed — or an inferred
one it should never have recorded — &lt;a href="https://github.com/Alberto-Codes/saucier/issues/new?template=extraction.yml"&gt;the issue template&lt;/a&gt;
asks for the entry number and the source lines, because a claim about this
project should be checkable the same way the project's own claims are.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;The release:&lt;/strong&gt; &lt;a href="https://github.com/Alberto-Codes/saucier/releases/tag/v0.2.0"&gt;saucier v0.2.0&lt;/a&gt;
— the resolver this post argues for: a parent may be any catalogued
preparation, derived goes from 29 to 50, three derivations the old rule
recorded are given back, and Bordelaise stays unresolved on purpose. MIT.&lt;/p&gt;
</content:encoded>
    </item>
    <item>
      <title>There is no model in this parser. It still told me ice cream was a sauce.</title>
      <link>https://alberto.codes/blog/2026-08-19-there-is-no-model-in-this-parser</link>
      <guid isPermaLink="true">https://alberto.codes/blog/2026-08-19-there-is-no-model-in-this-parser</guid>
      <pubDate>Wed, 19 Aug 2026 00:00:00 GMT</pubDate>
      <description>Post one of the saucier series reads a century-old cookbook with a regular expression — 124 sauces, 29 lineages, 95 that name no base. The first version of that census said 166 and 64, and forty of those were soups, jam, and a vanilla ice cream. Determinism did not catch that. Line numbers did.</description>
      <content:encoded>&lt;p&gt;Here is something I wanted to know. It is 2026, every conversation that starts
with &amp;quot;I have a pile of documents&amp;quot; ends with which model to point at them, and I
had a century-old cookbook sitting on my disk. How much of it can a regular
expression read?&lt;/p&gt;
&lt;p&gt;Not as a stunt, and not as an argument against models — later posts hand the
same book to one and find out what it does better. Just as a question worth
answering first, because it costs an afternoon.&lt;/p&gt;
&lt;p&gt;So: no model, no API key, no GPU, no network. &lt;code&gt;git clone&lt;/code&gt;,
&lt;code&gt;git checkout v0.1.0&lt;/code&gt;, &lt;code&gt;uv sync&lt;/code&gt;, &lt;code&gt;uv run saucier parse&lt;/code&gt;, and you are
standing where I am standing. Every post in this series adds one ingredient
and everything runs at every tag — a sauce should taste like something at
every stage of its reduction, and this is the earliest stage there is. The
checkout matters: &lt;code&gt;main&lt;/code&gt; keeps moving, and the numbers in this post are the
numbers at &lt;code&gt;v0.1.0&lt;/code&gt;.&lt;/p&gt;
&lt;p&gt;The useful part of it is not what the parser got right. It is what it got
wrong, how badly, and what caught it — because the thing that caught it is the
one property I would keep if I had to throw away everything else in the repo.&lt;/p&gt;
&lt;h2&gt;Reading a book literally&lt;/h2&gt;
&lt;p&gt;Sauces specifically, rather than the whole book. A sauce is the most
process-dense thing in a kitchen — reductions, emulsions and suspensions, with
temperature and timing and order that genuinely matter — and Escoffier's
mothers-and-derivatives scheme is already a taxonomy of them. The structure is
in the book. The only question is how much of it comes back out.&lt;/p&gt;
&lt;p&gt;Quite a lot, because Escoffier structured his own work. Every preparation in
&lt;em&gt;A Guide to Modern Cookery&lt;/em&gt; is numbered and titled — 2,963 of them — and he
names his five base sauces in a sentence you can point at:&lt;/p&gt;
&lt;blockquote&gt;
&lt;ol start="7"&gt;
&lt;li&gt;The basic sauces: Espagnole, Velouté, Béchamel, Tomato, and Hollandaise.&lt;/li&gt;
&lt;/ol&gt;
&lt;/blockquote&gt;
&lt;p&gt;So the five mothers are not in the source code. A regular expression reads them
out of the book:&lt;/p&gt;
&lt;pre&gt;&lt;code class="language-python"&gt;MOTHERS = re.compile(r&amp;quot;basic sauces?:\s*(.+?)\.&amp;quot;, re.IGNORECASE | re.DOTALL)
&amp;quot;&amp;quot;&amp;quot;The source names its own base preparations; it is not our place to guess them.&amp;quot;&amp;quot;&amp;quot;
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Hardcoding those five strings produces identical output today and a lie the
first time the same code meets Fannie Farmer. Run the whole thing and it
reports:&lt;/p&gt;
&lt;pre&gt;&lt;code class="language-console"&gt;$ uv run saucier parse
source      escoffier-1907
mothers     bechamel, espagnole, hollandaise, tomato, veloute
sauces      124
derived     29 linked to a mother
unresolved  95 state no base in their prose
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;&lt;em&gt;Correction, 2026-09-01: that &lt;code&gt;escoffier-1907&lt;/code&gt; is the identifier this release
printed, and it was wrong. The file is the New and Revised Edition of January
1909, which says so on its own title page. The counts, entries and line
numbers here are unaffected and still reproduce at &lt;code&gt;v0.1.0&lt;/code&gt;.
&lt;a href="https://alberto.codes/blog/2026-09-01-i-added-a-second-copy-of-the-same-book"&gt;Post three&lt;/a&gt; is the
correction.&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;29 of 124 is a coverage number that would embarrass most extraction demos. Hold
that thought.&lt;/p&gt;
&lt;h2&gt;The first census was worse than embarrassing. It was wrong.&lt;/h2&gt;
&lt;p&gt;Before those numbers, this project published different ones: 166 sauces, 64
derivations. They were in the README, on the landing page, in the CLI output,
and in the argument for why a model comes later. They were also nonsense.&lt;/p&gt;
&lt;p&gt;An entry qualified as a sauce if its heading said &amp;quot;sauce&amp;quot;, or if the folded
heading contained a mother concept anywhere inside it. That second rule is a
substring test, and two of the five mothers Escoffier names — &lt;em&gt;tomato&lt;/em&gt;,
&lt;em&gt;velouté&lt;/em&gt; — are ordinary words in a cookery book. The catalogue filled up
accordingly: 25 velouté &lt;strong&gt;soups&lt;/strong&gt; from the soup chapter, eight tomato dishes
and preserves including &lt;code&gt;TOMATO JAM&lt;/code&gt; and &lt;code&gt;TOMATO SALAD&lt;/code&gt;, six fish and meat
dishes including &lt;code&gt;SOLE A LA HOLLANDAISE&lt;/code&gt;, and &lt;code&gt;BOMBE HOLLANDAISE&lt;/code&gt;, which is
vanilla ice cream in a mould.&lt;/p&gt;
&lt;p&gt;Roughly forty of the 166 were not sauces. Thirty of the 64 recorded
derivations belonged to entries that are not sauces. And the record it
produced looked like this:&lt;/p&gt;
&lt;pre&gt;&lt;code class="language-json"&gt;{ &amp;quot;title&amp;quot;: &amp;quot;GRILLED TOMATOES&amp;quot;, &amp;quot;parent&amp;quot;: &amp;quot;tomato&amp;quot;,
  &amp;quot;ref&amp;quot;: { &amp;quot;entry&amp;quot;: 2263, &amp;quot;line&amp;quot;: 36212 } }
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;A grilled tomato, recorded as a derivative of the mother sauce Tomato, because
the word &amp;quot;tomatoes&amp;quot; occurs in its prose. An absence of evidence turned into a
derivation and published with a line number.&lt;/p&gt;
&lt;p&gt;There was a second defect underneath it. When an entry's opening paragraph
named two mothers, &lt;code&gt;resolve_parent&lt;/code&gt; sorted the candidates and took the first
alphabetically. It decided three entries that way and was wrong in all three.
&lt;code&gt;SHRIMP SAUCE&lt;/code&gt; says &amp;quot;fish velouté or, failing this, Béchamel&amp;quot; — the source
named both and chose neither, and the parser answered Béchamel because B sorts
before V. That is an arbitrary choice wearing the costume of determinism.&lt;/p&gt;
&lt;h2&gt;Determinism bought reproducibility, not accuracy&lt;/h2&gt;
&lt;p&gt;The argument for writing a parser before reaching for a model usually runs like
this: a model asked for JSON returns well-formed JSON, with a plausible parent
for every entry, and those outputs pass validation, read correctly, and are
wrong. Confidently wrong output that validates is the failure mode you cannot
see.&lt;/p&gt;
&lt;p&gt;Every word of that is true, and none of it is a property of models. My parser
did it. No weights, no sampling, no temperature. It produced a well-formed,
schema-valid, internally consistent catalogue in which vanilla ice cream was a
sauce and grilled tomatoes descended from a mother sauce. It did so
reproducibly, which meant only that it was wrong the same way every time.&lt;/p&gt;
&lt;p&gt;A deterministic system is not a system that cannot be wrong. It is a system
whose wrongness holds still while you look at it. That turns out to be worth a
great deal — but only if something makes you look.&lt;/p&gt;
&lt;h2&gt;Who caught it&lt;/h2&gt;
&lt;p&gt;Not a test. No reasonable test asserts that a cookbook parser hasn't found ice
cream.&lt;/p&gt;
&lt;p&gt;Copilot did, reviewing the pull request:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;code&gt;is_sauce&lt;/code&gt; treats any title containing a mother concept id as a sauce. In the
committed corpus this pulls in clear non-sauce entries such as &lt;code&gt;390—MOCK TOMATOES&lt;/code&gt; (folds to &lt;code&gt;mock-tomatoes&lt;/code&gt;, which matches &lt;code&gt;tomato&lt;/code&gt;) and dish
headings like &lt;code&gt;829—SOLE A LA HOLLANDAISE&lt;/code&gt;. This will inflate the catalogue
and break the documented/published counts.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;It flagged the alphabetical parent bug in the same review, and it was right
about both.&lt;/p&gt;
&lt;p&gt;Two separate properties of this release made that review possible, and they are
worth pulling apart, because only one of them is the one people skip.&lt;/p&gt;
&lt;p&gt;The first is that the output is legible. The catalogue is a JSON file of
titles a person can read, and &lt;code&gt;BOMBE HOLLANDAISE&lt;/code&gt; sitting in a list of sauces
is wrong on sight — no line number required. That is why this release writes
JSON files rather than a database: while a human is still verifying a parser,
inspectable by eye is the correct storage format.&lt;/p&gt;
&lt;p&gt;The second is provenance, and it is what turned a suspicion into a finding.
Every claim names an entry and a line in a file that ships with the repo:&lt;/p&gt;
&lt;pre&gt;&lt;code class="language-console"&gt;$ sed -n '1680p' corpus/escoffier-1907.txt
32—SAUCE BORDELAISE
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Notice what the review quoted: &lt;code&gt;390—MOCK TOMATOES&lt;/code&gt;, &lt;code&gt;829—SOLE A LA HOLLANDAISE&lt;/code&gt;. Entry numbers, because entry numbers were there to quote. That is
the difference between &amp;quot;this rule looks too loose&amp;quot; and a report you can act on
in an afternoon — and it is what made the damage countable afterwards rather
than merely regrettable: forty of 166, thirty of 64.&lt;/p&gt;
&lt;p&gt;Neither property has anything to do with determinism. A model that emitted the
same catalogue with the same fields would have been exactly as catchable. One
that emitted bare strings with no references would not have been, and neither
would this parser.&lt;/p&gt;
&lt;p&gt;Four things in this release do that job, and they are the four I would keep if
I started over on a different book tomorrow:&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Commitment&lt;/th&gt;
&lt;th&gt;Costs&lt;/th&gt;
&lt;th&gt;Buys&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Every claim carries source id, entry, and line&lt;/td&gt;
&lt;td&gt;Two integers per record&lt;/td&gt;
&lt;td&gt;A wrong answer becomes a findable one&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Terms carry a language tag, never a translation&lt;/td&gt;
&lt;td&gt;One field&lt;/td&gt;
&lt;td&gt;&lt;em&gt;Nixtamal&lt;/em&gt; survives as &lt;em&gt;nixtamal&lt;/em&gt;, not &amp;quot;corn&amp;quot;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Unresolved is recorded as unresolved, never as &amp;quot;none&amp;quot;&lt;/td&gt;
&lt;td&gt;A worse-looking coverage number&lt;/td&gt;
&lt;td&gt;A later stage can fill it without overwriting a fact&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Ambiguity resolves to nothing&lt;/td&gt;
&lt;td&gt;Three fewer derivations&lt;/td&gt;
&lt;td&gt;No arbitrary choice is ever recorded as a reading&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;None of those require a model to be absent. They are what makes a model
&lt;em&gt;addable&lt;/em&gt; later without destroying the record — extraction sits behind a port,
so the thing that fills &lt;code&gt;parent&lt;/code&gt; can change while every guarantee about
traceability holds.&lt;/p&gt;
&lt;h2&gt;Who gets to decide what a sauce is&lt;/h2&gt;
&lt;p&gt;The fix is not a better substring test. It is a rule about authority.&lt;/p&gt;
&lt;p&gt;An entry now enters the catalogue on evidence the source supplies. Its heading
uses the singular word &amp;quot;sauce&amp;quot; before any &amp;quot;with&amp;quot; — so &lt;code&gt;SOUBISE SAUCE WITH RICE&lt;/code&gt;
is a sauce served with something, and &lt;code&gt;ASPARAGUS WITH VARIOUS SAUCES&lt;/code&gt; is
something served with a sauce. Or its heading names a mother and Escoffier
filed it in one of the three chapters he titled &lt;code&gt;THE LEADING WARM SAUCES&lt;/code&gt;, &lt;code&gt;THE SMALL COMPOUND SAUCES&lt;/code&gt;, and &lt;code&gt;COLD SAUCES AND COMPOUND BUTTERS&lt;/code&gt;.&lt;/p&gt;
&lt;p&gt;Reading those chapter titles is the same move as reading the mothers out of the
text. Deciding for ourselves that a velouté soup is not a sauce would not be,
however obviously true it is. The distinction sounds pedantic until you notice
that the first rule was exactly that kind of self-authored judgment, and it
put ice cream in a sauce catalogue.&lt;/p&gt;
&lt;p&gt;That is &lt;a href="https://github.com/Alberto-Codes/saucier/blob/v0.1.0/docs/adr/0007-the-source-classifies-its-own-contents.md"&gt;ADR-0007&lt;/a&gt;.
The census now lives in one place that the tests and the documentation both
read, so the next time a number moves it cannot move in only three of the four
places that publish it.&lt;/p&gt;
&lt;h2&gt;What is left: 95&lt;/h2&gt;
&lt;p&gt;Which brings back the number that was the point of all this. 95 preparations
name no mother where the parser looks.&lt;/p&gt;
&lt;p&gt;Not all of them are silent. Twenty-seven name another catalogued sauce in
their opening paragraph — Allemande, Normande, Bordelaise, Bercy, Madeira,
plain Butter Sauce — just not one of the five. Entry 45, Marrow Sauce, opens:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Follow the proportions as indicated under &amp;quot;Sauce Bordelaise&amp;quot; (No. 32) for the
necessary quantity of this sauce, the Marrow Sauce being only a variety of
the Bordelaise.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;An author stating a derivation, in plain English, with a cross-reference
number. &lt;code&gt;parent: null&lt;/code&gt;, because Bordelaise is not a mother. Escoffier's
structure is not a five-way star; it is a graph with depth, and resolving to
any catalogued sauce and walking the chain is a legitimate next rule —
deterministic, checkable line by line, no model required.&lt;/p&gt;
&lt;p&gt;&lt;img src="https://alberto.codes/saucier-unstated-chain.svg" alt="A chain of four sauces descending. At the top, Espagnole, entry 22 line 1392, the only one the source explicitly calls a mother. A dashed arrow marked &amp;quot;assumed, not stated&amp;quot; runs down to half-glaze, a reduction of Espagnole the book never spells out. Another dashed &amp;quot;assumed, not stated&amp;quot; arrow runs to Sauce Bordelaise, entry 32 line 1680, whose opening says &amp;quot;half-glaze&amp;quot; and never &amp;quot;Espagnole&amp;quot;, so its parent is null. A solid arrow marked &amp;quot;stated outright, still unresolved&amp;quot; runs to Marrow Sauce, entry 45 line 1895, which calls itself &amp;quot;only a variety of the Bordelaise&amp;quot; and whose parent is also null." /&gt;&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Correction, 2026-09-03: this figure and the paragraph beneath it say the book never spells out that half glaze is a reduction of Espagnole. It does, at entry 23, line 1437, in the first sentence of an entry the parser at this tag could not see. The figure is left as printed, because it shows what the parser recorded. &lt;a href="https://alberto.codes/blog/2026-09-03-the-book-spells-it-out-at-line-1437"&gt;The post about line 1437&lt;/a&gt; is the correction.&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;Which leaves roughly 68 that genuinely say nothing — and they are the
interesting ones. Bordelaise is among them. Its opening specifies shallots, red
wine, mignonette pepper, thyme, bay, and half a pint of half-glaze, and never
uses the word Espagnole, because a &lt;em&gt;demi-glace&lt;/em&gt; is an Espagnole reduction and
the reader knew that. The derivation is encoded in a term rather than a
sentence.&lt;/p&gt;
&lt;p&gt;Recovering knowledge an author assumed is exactly the work a language model can
do and a regular expression cannot. That is the case for adding one, and it is
now a priced case rather than an assertion: a model here has to beat 29,
against a source where I can tell you precisely which 68 entries it would have
to read correctly and which line each one is on.&lt;/p&gt;
&lt;h2&gt;What breaks next&lt;/h2&gt;
&lt;p&gt;A JSON file. That is the entire storage layer — parse the book, rewrite the
file, read it back to print a tree. It stops working the first time appending
one record means rewriting 124, or a question needs answering without loading
all of it. The next post is that failure and what replaces it, which is the
only way I am willing to introduce a database.&lt;/p&gt;
&lt;p&gt;Then one rung per post. A process graph instead of a flat record. NLP over the
68. A local model, then a frontier one, with the cost difference measured
rather than assumed. Video, where this schema has to survive input that cannot
be checked by eye — which is the real reason text came first, because text is
where you can still tell a schema bug from an extraction bug by reading.&lt;/p&gt;
&lt;p&gt;Every one of those replaces something in this release. The regex gets replaced.
The line numbers do not.&lt;/p&gt;
&lt;p&gt;None of which means a regular expression is the answer to your pile of
documents. It worked here because Escoffier numbered his entries, titled his
chapters, and named his own mothers — and even then it found ice cream. A book
that does none of that gives a parser nothing to read, and the first move there
is a different one.&lt;/p&gt;
&lt;p&gt;The repo is &lt;a href="https://github.com/Alberto-Codes/saucier"&gt;Alberto-Codes/saucier&lt;/a&gt;
and the parse takes about a second. The
&lt;a href="https://alberto-codes.github.io/saucier/"&gt;documentation site&lt;/a&gt; carries the
tutorial, the data model, the glossary, and the seven decision records — the
API reference on it is generated from the docstrings, so it cannot drift from
the code.&lt;/p&gt;
&lt;p&gt;If you find an entry the parser should have linked — or another ice cream —
there is
&lt;a href="https://github.com/Alberto-Codes/saucier/issues/new?template=extraction.yml"&gt;an issue template for exactly that&lt;/a&gt;.
It asks for the entry number and the source lines, because a claim about this
project should be checkable the same way the project's own claims are.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;The release:&lt;/strong&gt; &lt;a href="https://github.com/Alberto-Codes/saucier/releases/tag/v0.1.0"&gt;saucier v0.1.0&lt;/a&gt;
— the parser exactly as this post leaves it: 124 sauces, 29 linked to a mother,
95 that state no base in their prose, and every one of those claims citing a
line you can open with &lt;code&gt;sed&lt;/code&gt;. MIT.&lt;/p&gt;
</content:encoded>
    </item>
    <item>
      <title>Fit a model to the GPU you actually have</title>
      <link>https://alberto.codes/blog/2026-08-15-fit-a-model-to-the-gpu-you-actually-have</link>
      <guid isPermaLink="true">https://alberto.codes/blog/2026-08-15-fit-a-model-to-the-gpu-you-actually-have</guid>
      <pubDate>Sat, 15 Aug 2026 00:00:00 GMT</pubDate>
      <description>Measure a model's per-layer quantization damage, solve for a recipe that fits your card, and check the result before you trust it.</description>
      <content:encoded>&lt;h2&gt;Who this is for&lt;/h2&gt;
&lt;p&gt;You have a GPU with a fixed amount of memory. You want to run a model that
does not fit. You have already tried an off-the-shelf quantization and want
one fitted to your card instead of to the average card.&lt;/p&gt;
&lt;p&gt;Assumed: comfortable with a terminal, a CUDA GPU. Not assumed: any
quantization background — that is the
&lt;a href="https://alberto.codes/blog/2026-08-11-i-couldnt-tell-my-quantized-model-from-the-baseline"&gt;explanation post&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;Every command below was run start to finish on a clean rented 4090 with
nothing else on it. The output is pasted as it came back.&lt;/p&gt;
&lt;p&gt;&lt;img src="https://alberto.codes/vramfit-pipeline.svg" alt="Four steps. Scan produces a per-layer price list, plan turns it into a recipe under a memory ceiling, validate checks the real damage against the prediction and sends failures back to the solver, and pack builds the file." /&gt;&lt;/p&gt;
&lt;h2&gt;Before you start&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Install.&lt;/strong&gt; The base install is small on purpose, so the planning step runs
on a laptop with no GPU. The heavy parts are extras
(&lt;a href="https://github.com/Alberto-Codes/vramfit/blob/main/docs/adr/0005-heavy-deps-as-extras.md"&gt;ADR-0005&lt;/a&gt;):&lt;/p&gt;
&lt;pre&gt;&lt;code class="language-bash"&gt;pip install vramfit           # plan and budget only, no torch
pip install &amp;quot;vramfit[scan]&amp;quot;   # adds torch + transformers, for scan and validate
pip install &amp;quot;vramfit[pack]&amp;quot;   # adds what llama.cpp's converter needs
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;That base install is nine packages — vramfit, &lt;code&gt;typer&lt;/code&gt;, &lt;code&gt;structlog&lt;/code&gt;, and
typer's own dependencies:&lt;/p&gt;
&lt;pre&gt;&lt;code class="language-text"&gt;Pygments, annotated-doc, markdown-it-py, mdurl, rich,
shellingham, structlog, typer, vramfit

&amp;gt;&amp;gt;&amp;gt; import torch
ModuleNotFoundError: No module named 'torch'
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Two things that catch people. &lt;code&gt;vramfit validate&lt;/code&gt; needs &lt;code&gt;[scan]&lt;/code&gt;, not just
&lt;code&gt;scan&lt;/code&gt; — it replays a whole recipe through the same meter. And &lt;code&gt;[pack]&lt;/code&gt; does
&lt;strong&gt;not&lt;/strong&gt; give you llama.cpp; it provisions the interpreter its converter script
needs. You build llama.cpp yourself and point at the checkout with
&lt;code&gt;--llama-cpp&lt;/code&gt; (&lt;a href="https://github.com/Alberto-Codes/vramfit/blob/main/docs/adr/0012-gguf-type-mapping.md"&gt;ADR-0012&lt;/a&gt;).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Check that torch can actually see your card before anything else.&lt;/strong&gt; This
is the first thing that bites, and it isn't vramfit's doing. &lt;code&gt;pip&lt;/code&gt; resolves
a torch build for whatever CUDA it feels like, and on a clean 4090 box with
a 12.8 driver I got a 13.0 build:&lt;/p&gt;
&lt;pre&gt;&lt;code class="language-text"&gt;$ python -c &amp;quot;import torch; print(torch.__version__, torch.cuda.is_available())&amp;quot;
2.13.0+cu130 False
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;&lt;code&gt;nvidia-smi&lt;/code&gt; says &lt;code&gt;CUDA Version: 12.8&lt;/code&gt;, torch was built for 13.0, and so the
GPU may as well not exist. The fix is to name the index that matches your
driver:&lt;/p&gt;
&lt;pre&gt;&lt;code class="language-bash"&gt;pip install --force-reinstall torch --index-url https://download.pytorch.org/whl/cu128
&lt;/code&gt;&lt;/pre&gt;
&lt;pre&gt;&lt;code class="language-text"&gt;$ python -c &amp;quot;import torch; print(torch.__version__, torch.cuda.is_available())&amp;quot;
2.11.0+cu128 True
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Sixty seconds of checking against half an hour of a scan silently running on
CPU.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Disk.&lt;/strong&gt; More than you expect, because the pipeline works through an
uncompressed intermediate — an f16 GGUF that everything quantizes from.&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;3B walkthrough&lt;/th&gt;
&lt;th&gt;49B run&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;model weights&lt;/td&gt;
&lt;td&gt;5.8 GiB&lt;/td&gt;
&lt;td&gt;92 GiB&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;f16 GGUF&lt;/td&gt;
&lt;td&gt;5.75 GiB&lt;/td&gt;
&lt;td&gt;92.89 GiB&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;packed output&lt;/td&gt;
&lt;td&gt;2.41 GiB&lt;/td&gt;
&lt;td&gt;20.36 GiB&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;llama.cpp build&lt;/td&gt;
&lt;td&gt;~2 GiB&lt;/td&gt;
&lt;td&gt;~2 GiB&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;Budget roughly 20 GiB free for a 3B and 115 GiB for a 49B. The f16 conversion
is the step that surprises people, not the model download.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Time.&lt;/strong&gt; Measured, not estimated:&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;step&lt;/th&gt;
&lt;th&gt;3B on a 4090&lt;/th&gt;
&lt;th&gt;49B on a 4090&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;scan&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;52 min&lt;/strong&gt; (148 cells)&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;37.0 hours&lt;/strong&gt; (328 cells)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;plan&lt;/td&gt;
&lt;td&gt;seconds&lt;/td&gt;
&lt;td&gt;seconds&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;f16 convert&lt;/td&gt;
&lt;td&gt;2 min 12 s&lt;/td&gt;
&lt;td&gt;~1 hour&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;quantize&lt;/td&gt;
&lt;td&gt;24 s&lt;/td&gt;
&lt;td&gt;16 min 34 s&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;smoke test&lt;/td&gt;
&lt;td&gt;seconds&lt;/td&gt;
&lt;td&gt;seconds&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;The 49B scan streams weights from host RAM, which is most of why it is 40
times slower per cell than a model that fits resident.&lt;/p&gt;
&lt;h2&gt;1. Work out your real budget&lt;/h2&gt;
&lt;p&gt;The number that matters is not your card's size. It is your card's size
minus what the KV cache will need at the context length you plan to serve.
&lt;code&gt;vramfit budget&lt;/code&gt; does that arithmetic from the model's own &lt;code&gt;config.json&lt;/code&gt;,
which matters here because this model's attention blocks are NAS-pruned and
heterogeneous — you can't eyeball the cost per token from the parameter
count.&lt;/p&gt;
&lt;pre&gt;&lt;code class="language-console"&gt;$ vramfit budget --vram 24GiB --context 16384 --model-config config.json
attention layers      49  (KV 200704 bytes/token, fp16)
VRAM total            24.00 GiB
- KV cache            3.06 GiB  (16384 tokens x 1 seq)
- runtime overhead    2.00 GiB
= weight budget       18.94 GiB
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;That remainder is what the solver targets, to the byte.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Do this before anything else, because it can end the job right here.&lt;/strong&gt; Solving is
cheap — seconds, no GPU — so you can find out whether your ceiling is even
reachable before spending 37 hours on a scan. Run against a published price
list and the answer comes back immediately:&lt;/p&gt;
&lt;pre&gt;&lt;code class="language-text"&gt;error: no recipe fits the 12.47 GiB weight budget
       — minimum achievable is 15.31 GiB (2.85 GiB over)
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;That is a 16 GiB card being told this model will not fit, at any precision
this price list allows. Better to find out now than after the download.&lt;/p&gt;
&lt;p&gt;If your budget does land on a stock preset's size, take the preset — this
whole pipeline buys you the most when your ceiling sits &lt;em&gt;between&lt;/em&gt; the sizes
the shelf stocks. The companion post works three ceilings side by side, including where the
recipe's priorities reorder rather than just shrink.&lt;/p&gt;
&lt;h2&gt;2. Scan&lt;/h2&gt;
&lt;p&gt;Quantize one layer group at a time, measure how far the output distribution
moves, write the per-layer price list. This is the expensive step, and it
produces the artifact worth keeping.&lt;/p&gt;
&lt;p&gt;You need a calibration text — a few hundred kilobytes of prose the meter
pushes through the model to see what moves. Any representative text works.
For a run you can compare against mine, take the one published beside the
49B maps:&lt;/p&gt;
&lt;pre&gt;&lt;code class="language-bash"&gt;curl -LO https://huggingface.co/datasets/Alberto-Codes/Llama-3_3-Nemotron-Super-49B-v1_5-sensitivity-maps/resolve/main/calibration.txt
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;That file is 754 KiB of WikiText-2, and it's the same one every number on
this page was measured with.&lt;/p&gt;
&lt;p&gt;Here it is on Qwen2.5-3B-Instruct, which is small enough to finish while you
have lunch:&lt;/p&gt;
&lt;pre&gt;&lt;code class="language-bash"&gt;PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True \
vramfit scan /path/to/Qwen2.5-3B-Instruct \
  --calibration calibration.txt \
  --max-tokens 32768 \
  --precisions 8,4,3,2 \
  --group-by layer \
  --out sensitivity.json
&lt;/code&gt;&lt;/pre&gt;
&lt;pre&gt;&lt;code class="language-text"&gt;[3/148] model.embed_tokens @ 3-bit damage 0.249366
[8/148] model.layers.0 @ 2-bit damage 0.905969
[20/148] model.layers.3 @ 2-bit damage 2.908759
...
[147/148] model.layers.35 @ 3-bit damage 0.077071
scanned 37 groups x 4 precisions over 32768 tokens -&amp;gt; sensitivity.json
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;It prints a line per cell, so you can tell it's alive. What you want to see
is 8-bit damage in the thousandths and 2-bit damage one to three orders
higher — that spread is the signal. If every precision costs about the same,
your calibration text is too short to have converged.&lt;/p&gt;
&lt;p&gt;37 groups, 4 precisions, 148 cells, 52 minutes. Look at what it found before
you go further: &lt;code&gt;model.layers.3&lt;/code&gt; costs 2.9088 at 2-bit and 0.0011 at 8-bit,
a factor of about 2,600 between the cheapest and dearest thing you can do to
one layer. That spread is the entire reason this tool exists.&lt;/p&gt;
&lt;p&gt;The map is a property of the model, not of your machine. Scan once and you
can solve it against any ceiling later, on a laptop, without a GPU — which
is what makes step 3 cheap.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;The memory cap matters on bigger models.&lt;/strong&gt; The 3B above fits a 4090
resident and needs no cap at all. The 49B needs &lt;code&gt;--gpu-memory 15GiB&lt;/code&gt; on the
same card, and dies at 17:&lt;/p&gt;
&lt;pre&gt;&lt;code class="language-text"&gt;error: scan halted at model.embed_tokens 8-bit: CUDA out of memory.
Tried to allocate 1.96 GiB ... (checkpoint keeps 0 cells)
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;The cap is lower than the card because the meter needs headroom &lt;em&gt;beside&lt;/em&gt; the
weights for the reference activations. &lt;code&gt;expandable_segments&lt;/code&gt; is in the
command above for the same reason — without it, fragmentation kills a cell
that would otherwise fit. A halted scan checkpoints, so &lt;code&gt;--resume&lt;/code&gt; picks up
where it stopped.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Calibration length is the setting people get wrong.&lt;/strong&gt; 8k tokens is a pilot
— not a scan. Re-planning the &lt;em&gt;same budget&lt;/em&gt; on a 32k map instead of an 8k one
flipped &lt;strong&gt;41 of 82 assignments&lt;/strong&gt; on the 49B, and predicted damage went from
0.4949 to 0.0940. 32k suffices at 3-bit and above. (The 3B walkthrough above
used 32k for the same reason.)&lt;/p&gt;
&lt;h2&gt;3. Plan&lt;/h2&gt;
&lt;p&gt;Solve the price list against your ceiling. Seconds, no GPU.&lt;/p&gt;
&lt;pre&gt;&lt;code class="language-bash"&gt;vramfit plan sensitivity.json --vram 4GiB --kv-headroom 1.5625GiB --out recipe.json
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;The thing worth understanding is that this does nothing interesting until
your ceiling actually squeezes. Same 3B map, four cards:&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;card&lt;/th&gt;
&lt;th&gt;weight budget&lt;/th&gt;
&lt;th&gt;result&lt;/th&gt;
&lt;th&gt;downgrades&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;12 GiB&lt;/td&gt;
&lt;td&gt;9.44 GiB&lt;/td&gt;
&lt;td&gt;3.07 GiB, every group at 8-bit&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;6 GiB&lt;/td&gt;
&lt;td&gt;3.44 GiB&lt;/td&gt;
&lt;td&gt;3.07 GiB, every group at 8-bit&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;4 GiB&lt;/td&gt;
&lt;td&gt;2.44 GiB&lt;/td&gt;
&lt;td&gt;2.42 GiB, 19 at 8-bit and 18 at 4-bit&lt;/td&gt;
&lt;td&gt;18&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;3 GiB&lt;/td&gt;
&lt;td&gt;1.44 GiB&lt;/td&gt;
&lt;td&gt;1.43 GiB, 17 at 4-bit and 20 at 3-bit&lt;/td&gt;
&lt;td&gt;57&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;A 3B on a 12 GiB card is not a problem anyone has. Run this on a model that
genuinely doesn't fit, or you'll conclude the solver does nothing — it just
has nothing to decide. Predicted damage across those four rows runs 0.0340,
0.0340, 0.0861, 0.4973.&lt;/p&gt;
&lt;p&gt;Two flags deserve plain explanations, because they're what closed the gap on
the published 49B artifact.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;&lt;code&gt;--protect &amp;quot;glob=bits&amp;quot;&lt;/code&gt;&lt;/strong&gt; holds specific &lt;em&gt;tensors&lt;/em&gt; at a precision floor
inside their group
(&lt;a href="https://github.com/Alberto-Codes/vramfit/blob/main/docs/adr/0022-within-layer-protections.md"&gt;ADR-0022&lt;/a&gt;).
A layer group is one assignment, but a layer isn't uniform — the attention
value projection can be the thing that breaks while the rest of the layer is
fine at 3-bit. Protection buys that one tensor back without paying for the
whole group. The published recipe carries 48 such pairs.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;&lt;code&gt;--exclude-imatrix &amp;quot;glob&amp;quot;&lt;/code&gt;&lt;/strong&gt; quantizes a matched protected tensor &lt;em&gt;without&lt;/em&gt;
its importance-matrix rows
(&lt;a href="https://github.com/Alberto-Codes/vramfit/blob/main/docs/adr/0023-imatrix-exclusions.md"&gt;ADR-0023&lt;/a&gt;).
It sounds backwards, and it's a remedy rather than a default: occasionally
the importance weighting makes a specific tensor's fit collapse, and the
reconstruction check below is what tells you which one.&lt;/p&gt;
&lt;h2&gt;4. Pack&lt;/h2&gt;
&lt;pre&gt;&lt;code class="language-bash"&gt;vramfit pack recipe.json \
  --model /path/to/Qwen2.5-3B-Instruct \
  --llama-cpp /path/to/llama.cpp \
  --smoke-text smoke.txt \
  --out qwen3b-fit4gib.gguf
&lt;/code&gt;&lt;/pre&gt;
&lt;pre&gt;&lt;code class="language-text"&gt;converting /path/to/Qwen2.5-3B-Instruct -&amp;gt; qwen3b-f16.gguf (minutes at 3B scale)
packed 37 groups -&amp;gt; qwen3b-fit4gib.gguf (2.41 GiB), weight budget 2.44 GiB, margin 24.23 MiB under
smoke test: perplexity 16.0119 over 2 chunks, ceiling 1000 — passed
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;The run log records what actually happened, which is what you want six
months later:&lt;/p&gt;
&lt;pre&gt;&lt;code class="language-text"&gt;[gguf_converted] bytes=6178317216, reused=False, seconds=131.529
[model_packed]   base_type=Q4_K_S, output_tensor_type=q8_0, overrides=36, seconds=23.705
[size_checked]   fits=True, margin_bytes=25402464, weight_budget_bytes=2617245696
[smoke_tested]   chunks=2, passed=True, perplexity=16.0119, threshold=1000.0
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Note &lt;code&gt;base_type&lt;/code&gt;. The recipe drives every override, but the floor still maps
to a stock ftype, so a pack that goes wrong tends to go wrong by quietly
falling back to that floor rather than by failing.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;On bigger models the importance matrix stops being optional.&lt;/strong&gt; The 3B pack
above ran without one — &lt;code&gt;imatrix=None&lt;/code&gt; in that log. At 3-bit on the 49B it
was the single largest contributor to closing the gap with the
baseline — the scoreboard puts it at about &lt;strong&gt;81 % of the recipe's deficit&lt;/strong&gt;.
Pass it with &lt;code&gt;--imatrix&lt;/code&gt;, which also lets pack run the per-tensor
reconstruction check below.&lt;/p&gt;
&lt;h2&gt;5. Check it before you trust it&lt;/h2&gt;
&lt;p&gt;Non-negotiable, and the reason is a real incident: a recipe once predicted
damage 1.44 and the packed artifact was &lt;strong&gt;destroyed&lt;/strong&gt; — perplexity around
10⁶, top-token agreement 0.3 %. Nothing between plan and pack would have
caught it.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;The smoke test&lt;/strong&gt; (&lt;code&gt;--smoke-text&lt;/code&gt;) runs a short perplexity pass on the
packed file and fails above a ceiling
(&lt;a href="https://github.com/Alberto-Codes/vramfit/blob/main/docs/adr/0017-post-pack-smoke-test.md"&gt;ADR-0017&lt;/a&gt;).
The 3B run above scored 16.0119 against a ceiling of 1000.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;That number is not a quality claim, and you shouldn't read it as one.&lt;/strong&gt;
The smoke test runs two chunks on the packed file and nothing else. It has
no f16 reference to compare against, so it can tell you the artifact still
produces language — it cannot tell you how much you lost. The ceiling looks
absurdly loose because it's a corpse detector: the 10⁶ artifact would have
tripped it instantly, and that is the whole job. Measuring what you actually
lost is &lt;code&gt;vramfit validate&lt;/code&gt; and a real evaluation, which is the subject of
&lt;a href="https://alberto.codes/blog/2026-08-11-i-couldnt-tell-my-quantized-model-from-the-baseline"&gt;the last post&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;The reconstruction check&lt;/strong&gt; runs when you pack with &lt;code&gt;--imatrix&lt;/code&gt; and
protections. It re-quantizes each protected tensor and measures how far it
moved, naming any tensor whose fit collapsed
(&lt;a href="https://github.com/Alberto-Codes/vramfit/blob/main/docs/adr/0022-within-layer-protections.md"&gt;ADR-0022&lt;/a&gt;).
On the 49B it took 4 min 36 s and passed. When it doesn't pass, the tensor it
names is the input to &lt;code&gt;--exclude-imatrix&lt;/code&gt; from step 3.&lt;/p&gt;
&lt;p&gt;Neither check tells you the artifact is &lt;em&gt;good&lt;/em&gt;. They tell you it isn't
broken, which is the more urgent question.&lt;/p&gt;
&lt;h2&gt;When this is the wrong tool&lt;/h2&gt;
&lt;p&gt;Off-the-shelf quantizations are good, they're free, and they're one download.
This is worth it when your budget is unusual, when you need to know the
artifact has no cliffs, or when you want the evidence.&lt;/p&gt;
&lt;p&gt;The four-ceiling table in step 3 is the sharpest version of that test. If
your ceiling leaves the model comfortable, the solver has nothing to decide
and you should take a preset. The further your real ceiling sits from the
sizes the shelf happens to stock, the more this pays — and
&lt;a href="https://alberto.codes/blog/2026-08-15-a-different-ceiling-is-a-different-recipe"&gt;the companion post&lt;/a&gt;
works that through on a model where the shelf runs out entirely.&lt;/p&gt;
</content:encoded>
    </item>
    <item>
      <title>A different ceiling is a different recipe. I finally checked.</title>
      <link>https://alberto.codes/blog/2026-08-15-a-different-ceiling-is-a-different-recipe</link>
      <guid isPermaLink="true">https://alberto.codes/blog/2026-08-15-a-different-ceiling-is-a-different-recipe</guid>
      <pubDate>Sat, 15 Aug 2026 00:00:00 GMT</pubDate>
      <description>I claimed a smaller card gives you a different quantization recipe, not the same one squeezed, and then didn't prove it. Here's that claim run against the published price list — including the ceiling where the honest answer is that there's no dish.</description>
      <content:encoded>&lt;p&gt;Cost a braise for a forty-dollar plate and you're refining: better stock,
finish the sauce properly. Cost the same braise for twenty-eight and you're
not shaving every ingredient by thirty percent. You drop the saffron
entirely and keep the technique — a thin version of the dish is worse than a
different dish done right.&lt;/p&gt;
&lt;p&gt;Cost it for fifteen and there's no version. The professional answer is to
say so before anyone shops.&lt;/p&gt;
&lt;p&gt;I wrote something in &lt;a href="https://alberto.codes/blog/2026-08-11-i-couldnt-tell-my-quantized-model-from-the-baseline"&gt;the last post&lt;/a&gt;
that I believed but hadn't checked:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;A 16 GiB card, or the same card serving twice the context, is a different
ceiling and therefore a different recipe — not the same recipe squeezed.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;Reasonable. Unproven. Only one ceiling had receipts, so I moved on. This is
me going back — and the sentence turns out to be too smooth. A different
ceiling isn't only a different recipe. Sometimes it's a question with no
answer at all.&lt;/p&gt;
&lt;h2&gt;Why this is cheap to check&lt;/h2&gt;
&lt;p&gt;The expensive half of measuring a model is the scan. Crush one layer group
at a time, measure how far the output moves, write down the price. On a 49B
that's 37 hours.&lt;/p&gt;
&lt;p&gt;But the scan doesn't know your card. It produces a &lt;strong&gt;price list&lt;/strong&gt; — what
each layer group costs you at each precision — and that list is the same
whatever you plan to run it on. This is mise en place. Solving that list
against a memory ceiling is a separate step, and it takes seconds on a
laptop with no GPU.&lt;/p&gt;
&lt;p&gt;So — one price list, three ceilings. The list is
&lt;a href="https://huggingface.co/datasets/Alberto-Codes/Llama-3_3-Nemotron-Super-49B-v1_5-sensitivity-maps"&gt;published&lt;/a&gt;,
so you can run this yourself.&lt;/p&gt;
&lt;h2&gt;Your card's size isn't your budget&lt;/h2&gt;
&lt;p&gt;What the solver targets is what's left after the model's conversation memory
is reserved. On this model that's genuinely hard to eyeball — the attention
blocks are NAS-pruned and heterogeneous, so you can't infer the cost per
token from the parameter count.&lt;/p&gt;
&lt;pre&gt;&lt;code class="language-console"&gt;$ vramfit budget --vram 24GiB --context 16384 --model-config config.json
attention layers      49  (KV 200704 bytes/token, fp16)
VRAM total            24.00 GiB
- KV cache            3.06 GiB  (16384 tokens x 1 seq)
- runtime overhead    2.00 GiB
= weight budget       18.94 GiB
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Three ceilings: the card I've got at the context I published against, a
smaller card, and my card serving twice the context.&lt;/p&gt;
&lt;p&gt;&lt;img src="https://alberto.codes/vramfit-three-ceilings.svg" alt="One 37-hour scan produces a price list. Solving it at three ceilings gives three different outcomes: a 24 GiB card at 16k context fits at 20.46 GiB, the same card at 32k context produces a different recipe at 17.13 GiB with 21 of 82 assignments changed, and a 16 GiB card produces no recipe at all because the smallest possible arrangement is 15.31 GiB against a 12.47 GiB budget." /&gt;&lt;/p&gt;
&lt;h2&gt;What came back&lt;/h2&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;A: 24 GiB, 16k&lt;/th&gt;
&lt;th&gt;B: 16 GiB, 16k&lt;/th&gt;
&lt;th&gt;C: 24 GiB, 32k&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;weight budget&lt;/td&gt;
&lt;td&gt;20.47 GiB&lt;/td&gt;
&lt;td&gt;12.47 GiB&lt;/td&gt;
&lt;td&gt;17.41 GiB&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;recipe&lt;/td&gt;
&lt;td&gt;20.46 GiB&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;none exists&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;17.13 GiB&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;predicted damage&lt;/td&gt;
&lt;td&gt;0.1215&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;0.2212&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;bits used&lt;/td&gt;
&lt;td&gt;56x2, 13x3, 8x4, 5x8&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;71x2, 4x3, 6x4, 1x8&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;Column A reproduces the recipe I shipped, to the fourth decimal of predicted
damage — that is the control. Without it the other two columns are just
numbers I generated.&lt;/p&gt;
&lt;h2&gt;The most useful column is the empty one&lt;/h2&gt;
&lt;p&gt;A 16 GiB card can't run this model. Not &amp;quot;runs badly.&amp;quot; There is no recipe:&lt;/p&gt;
&lt;pre&gt;&lt;code class="language-text"&gt;error: no recipe fits the 12.47 GiB weight budget
       — minimum achievable is 15.31 GiB (2.85 GiB over)
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;This is the part I didn't anticipate when I wrote the original line.&lt;/p&gt;
&lt;p&gt;A preset can't tell you this. Presets are fixed rules — you browse file
sizes, find one under your number, download it. It will load. Whether the
thing inside still works is something you discover later — subjectively, over
days of it being slightly off.&lt;/p&gt;
&lt;p&gt;The solver's answer is arithmetic. Here is the smallest arrangement this
price list permits, here is your budget, the first is bigger by 2.85 GiB. One
command, no download.&lt;/p&gt;
&lt;p&gt;Being told no quickly is most of what I want from a tool that costs 37 hours
when the answer is yes.&lt;/p&gt;
&lt;p&gt;If that's your card, you have three real options and the tool won't pick for
you. Run a smaller model — a 27B or 32B at these precisions fits 16 GiB
comfortably. Rent a bigger card for the work that needs this one. Or accept
partial offload and a much slower model, which llama.cpp will do and vramfit
doesn't plan for. What you shouldn't do is download a preset that fits and
assume the fit means it works.&lt;/p&gt;
&lt;h2&gt;What actually changed between A and C&lt;/h2&gt;
&lt;p&gt;Same card, twice the context. My claim was &amp;quot;a different recipe, not the same
recipe squeezed.&amp;quot; The data says both, and neither.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;21 of 82 assignments differ.&lt;/strong&gt; 61 are identical. So it is not a different
recipe. It is also not the same one.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;The level drops&lt;/strong&gt;, as you'd expect: 56 groups at 2-bit becomes 71.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;The shape changes&lt;/strong&gt;, which I didn't expect. Of the 13 groups holding
4-bit or better at 16k, only 7 keep it. The biggest mover is the embedding
table, which falls from &lt;strong&gt;8-bit to 3-bit&lt;/strong&gt;. It is the largest single object
on the card, and the solver sells it to keep the layers alive. Layers 76
through 78, the deepest three, drop from 4-bit to 2-bit together.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;There is the saffron. At the tighter ceiling the solver doesn't shave
everything down a notch — it re-decides what's worth protecting at all, and
the first thing it gives up is the most expensive item on the counter.&lt;/p&gt;
&lt;p&gt;I'd have guessed the opposite. I'd have guessed it protects the big table
and squeezes the many small layers, because that is what &amp;quot;make it smaller&amp;quot;
feels like it should mean. The measurement disagrees, and the measurement
has receipts.&lt;/p&gt;
&lt;h2&gt;What a preset user does at each ceiling&lt;/h2&gt;
&lt;p&gt;This is what decides whether any of it is worth an afternoon.&lt;/p&gt;
&lt;p&gt;At &lt;strong&gt;A&lt;/strong&gt;, presets are fine. Honestly. &lt;code&gt;Q3_K_S&lt;/code&gt; is 20.45 GiB against a
20.47 GiB budget — it fits with 20 MiB to spare, and the measured recipe
beats it by a margin you need instruments to see. That was the whole subject
of the last post.&lt;/p&gt;
&lt;p&gt;At &lt;strong&gt;C&lt;/strong&gt; the shelf runs out. Every published preset I measured for this model
overflows a 17.41 GiB budget. The smallest, &lt;code&gt;IQ3_XXS&lt;/code&gt;, is 18.18 GiB and
misses by 0.77 GiB. So you take the next one down and wear the gap — at these
sizes, most of a bit per weight.&lt;/p&gt;
&lt;p&gt;At &lt;strong&gt;B&lt;/strong&gt; nothing works — and you learn it in one command instead of after a
download.&lt;/p&gt;
&lt;p&gt;The pattern is simple enough: the further your real ceiling sits from the
sizes the shelf happens to stock, the more measuring pays. If your budget
lands on a preset, take the preset.&lt;/p&gt;
&lt;h2&gt;One caveat, and it's the honest kind&lt;/h2&gt;
&lt;p&gt;Column C only exists because that solve was allowed to use 2-bit.&lt;/p&gt;
&lt;p&gt;The artifact I published was planned on a price list with the 2-bit column
removed, under a rule that bars 2-bit until it's been priced in the runtime
rather than in the measurement frame
(&lt;a href="https://github.com/Alberto-Codes/vramfit/blob/main/docs/adr/0021-runtime-frame-measurement.md"&gt;ADR-0021&lt;/a&gt;
decision 4). On that list &lt;strong&gt;both&lt;/strong&gt; B and C are infeasible — the floor is
20.06 GiB, and 32k of context leaves 17.41 GiB.&lt;/p&gt;
&lt;p&gt;So under the policy the shipped artifact was actually built with, this model
serves 16k of context on a 24 GiB card and nothing else. Doubling the
context isn't a tuning exercise. It is a question about whether you trust
2-bit, and on a different model that question has since been measured and
answered no, at 4.1 times the reference perplexity.&lt;/p&gt;
&lt;p&gt;I'd rather show you the column and tell you what it costs than quietly plan
with a width I've argued against everywhere else.&lt;/p&gt;
&lt;h2&gt;What I'd say now&lt;/h2&gt;
&lt;p&gt;The original sentence wasn't wrong — it was underspecified in a way that hid
the good part.&lt;/p&gt;
&lt;p&gt;&amp;quot;A different ceiling is a different recipe&amp;quot; sounds smooth — turn the dial,
get a different answer. What happens is that a quarter of the assignments
move, the solver's priorities reorder, and at some point the dial runs out
of travel and the honest output is an error message.&lt;/p&gt;
&lt;p&gt;If I were writing that line again: &lt;strong&gt;a different ceiling is a different
question, and sometimes it has no answer.&lt;/strong&gt;&lt;/p&gt;
&lt;hr /&gt;
&lt;p&gt;&lt;em&gt;Part of a series on measuring quantization damage. Start with
&lt;a href="https://alberto.codes/blog/2026-08-11-i-couldnt-tell-my-quantized-model-from-the-baseline"&gt;I couldn't tell my quantized model from the baseline&lt;/a&gt;,
which is where the idea comes from. If you want to run this yourself,
&lt;a href="https://alberto.codes/blog/2026-08-15-fit-a-model-to-the-gpu-you-actually-have"&gt;the walkthrough&lt;/a&gt;
takes the whole pipeline from install to a packed file. The full scoreboard
behind these numbers is still coming.&lt;/em&gt;&lt;/p&gt;
</content:encoded>
    </item>
    <item>
      <title>I couldn't tell my quantized model from the baseline. The instruments could.</title>
      <link>https://alberto.codes/blog/2026-08-11-i-couldnt-tell-my-quantized-model-from-the-baseline</link>
      <guid isPermaLink="true">https://alberto.codes/blog/2026-08-11-i-couldnt-tell-my-quantized-model-from-the-baseline</guid>
      <pubDate>Tue, 11 Aug 2026 00:00:00 GMT</pubDate>
      <description>I shrank a 93 GB model onto a 24 GB card by measuring which layers survive being crushed, instead of guessing. Then I couldn't tell the result from the standard quant by talking to it — which turns out to be the whole point.</description>
      <content:encoded>&lt;p&gt;I asked two compressed copies of the same model fifteen questions each. Same questions, same settings, no randomness.&lt;/p&gt;
&lt;p&gt;They scored &lt;strong&gt;19 out of 25. Both of them.&lt;/strong&gt; And they didn't just tie — they &lt;em&gt;agreed&lt;/em&gt;. Both got the same four facts wrong. Both wrote the same function with the same bug in it. One file was mine, built by measuring the model layer by layer. The other was the standard off-the-shelf quantization at that size, &lt;a href="https://huggingface.co/bartowski/nvidia_Llama-3_3-Nemotron-Super-49B-v1_5-GGUF"&gt;bartowski's Q3_K_S&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;Chat with both and you'd find nothing, and conclude my work bought nothing.&lt;/p&gt;
&lt;p&gt;It bought something. You just can't get at it by chatting — and why not is worth more than the result.&lt;/p&gt;
&lt;h2&gt;The problem, if you've never met it&lt;/h2&gt;
&lt;p&gt;Nemotron Super 49B is about &lt;strong&gt;93 GB&lt;/strong&gt; at full precision. An RTX 4090, a very good consumer card, has &lt;strong&gt;24 GB&lt;/strong&gt;. The model has to fit entirely inside that to run well.&lt;/p&gt;
&lt;p&gt;So you compress it. The weights are billions of numbers stored at 16 bits each; store them at 4 bits, or 3, and the file shrinks. That's quantization, and it's why you can run useful models at home at all.&lt;/p&gt;
&lt;p&gt;The catch: crushing the numbers damages the model, and &lt;strong&gt;not all parts damage equally.&lt;/strong&gt; Most layers take 3 bits without complaining. A few — the value projections at the front of the stack, and the very first block — need more, and crushing those degrades the whole thing in ways the file size never warns you about.&lt;/p&gt;
&lt;p&gt;Here's what surprised me. Essentially every GGUF you can download picks precision from a small family of &lt;strong&gt;presets&lt;/strong&gt; — fixed rules that branch on a few architectural details, but never measure &lt;em&gt;this&lt;/em&gt; model's fragility or solve against &lt;em&gt;your&lt;/em&gt; memory ceiling. (Some other formats do measure. EXL2 runs a per-layer error pass, though it targets an average bit rate rather than a hard VRAM budget.)&lt;/p&gt;
&lt;p&gt;Those presets are your abuelita's recipe. A handful of this, cook it till it looks right. They genuinely work, they're the product of real craft, and the reasons they work were never written down — nobody checks the rule against the model in front of them, because the rule has always been good enough.&lt;/p&gt;
&lt;p&gt;I wanted the America's Test Kitchen version. Same dish, weighed in grams, tested in kitchens that aren't hers, with the steps that matter separated from the ones that don't. Not a better rule. &lt;strong&gt;No rule at all.&lt;/strong&gt;&lt;/p&gt;
&lt;h2&gt;What compression actually takes away&lt;/h2&gt;
&lt;p&gt;One idea about how these models work explains everything else. It's the only technical thing in this post.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;A language model doesn't produce a word. It produces a probability distribution over every possible next word.&lt;/strong&gt; Tens of thousands of candidates, each with a score. &amp;quot;The capital of France is ___&amp;quot; might come out as &lt;em&gt;Paris&lt;/em&gt; 94%, &lt;em&gt;the&lt;/em&gt; 2%, &lt;em&gt;a&lt;/em&gt; 1%, and a long tail of everything else. Only then does something pick one.&lt;/p&gt;
&lt;p&gt;&lt;img src="https://alberto.codes/vramfit-distribution-shoulders.svg" alt="Two probability curves over possible next words. Both peak on the same top word, so greedy decoding picks identically from either. Behind the peak the curves separate — that difference is what KL divergence measures." /&gt;&lt;/p&gt;
&lt;p&gt;Now the trap is obvious. My fifteen questions used greedy decoding — always take the highest-scoring word. That reads &lt;strong&gt;only the tip of the peak.&lt;/strong&gt; Two copies can agree on the top word often enough that fifteen questions cannot separate them, while the shape behind it has drifted.&lt;/p&gt;
&lt;p&gt;Conversation samples the peak. Damage lives in the shoulders.&lt;/p&gt;
&lt;p&gt;The instrument that sees the shoulders is &lt;strong&gt;KL divergence&lt;/strong&gt; — how much two probability shapes differ. Point it at the original and a compressed copy across hundreds of pages, and you get one honest number: how far did this copy drift from what it was made from?&lt;/p&gt;
&lt;p&gt;That's the number I optimize. Not vibes, and not a preset.&lt;/p&gt;
&lt;h2&gt;Measure, then solve&lt;/h2&gt;
&lt;p&gt;&lt;a href="https://github.com/Alberto-Codes/vramfit"&gt;vramfit&lt;/a&gt; crushes one layer group at a time to build a price list for this specific model, solves that against your memory ceiling, checks the real damage against its own prediction, and only then builds the file.&lt;/p&gt;
&lt;p&gt;&lt;img src="https://alberto.codes/vramfit-pipeline.svg" alt="Four steps. Scan produces a per-layer price list, plan turns it into a recipe under a memory ceiling, validate checks the real damage against the prediction and sends failures back to the solver, and pack builds the file." /&gt;&lt;/p&gt;
&lt;p&gt;The output isn't a preset. It's a recipe fitted to one model and one card — and the &amp;quot;one card&amp;quot; part is where this really differs.&lt;/p&gt;
&lt;p&gt;Your card's size is not your budget. A 24 GiB card doesn't give you 24 GiB for weights: the model also needs room for the conversation it's holding, and reserving 16k of context left me &lt;strong&gt;20.47 GiB&lt;/strong&gt;. That remainder is what the solver targets, to the byte.&lt;/p&gt;
&lt;p&gt;&lt;img src="https://alberto.codes/vramfit-budget.svg" alt="A 24 GiB card split into 20.47 GiB for weights and 3.53 GiB reserved for 16k of context. The remainder is what the solver targets, and a different card or a longer context moves the line." /&gt;&lt;/p&gt;
&lt;p&gt;A 16 GiB card, or the same card serving twice the context, is a different ceiling and therefore a different recipe — not the same recipe squeezed. With presets you browse fixed sizes and take whichever fits under your number, wearing whatever gap is left over. Here the number goes in the front.&lt;/p&gt;
&lt;p&gt;At the same file size, the measured recipe drifts &lt;strong&gt;less&lt;/strong&gt; from the original — a 2.9% smaller gap, tiny-sounding but far outside the measurement noise. It also beats the three i-quants that fit the same budget, and ties the baseline on five capability benchmarks, so the closer fit costs nothing the benchmarks can see.&lt;/p&gt;
&lt;p&gt;There's a sharper payoff than a better average. An earlier candidate of mine passed every check and looked clean — until a page-by-page comparison found one page out of 564 where it fell off a cliff, diverging &lt;strong&gt;more than fifty times&lt;/strong&gt; worse than the baseline on that text. One page, and nothing had flagged it. The published artifact is the first with no such page anywhere in the 564.&lt;/p&gt;
&lt;p&gt;That's the foolproof part, and the real return on weighing your ingredients. Not a better score — a &lt;strong&gt;known shape.&lt;/strong&gt;&lt;/p&gt;
&lt;h2&gt;The part nobody publishes&lt;/h2&gt;
&lt;p&gt;Here's what I actually think the edge is, and it isn't the 2.9%.&lt;/p&gt;
&lt;h3&gt;Five losses, and what each one bought&lt;/h3&gt;
&lt;p&gt;Before that recipe won, &lt;strong&gt;it lost five times.&lt;/strong&gt;&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Full-loop attempt&lt;/th&gt;
&lt;th&gt;Date&lt;/th&gt;
&lt;th&gt;Result&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;First real attempt&lt;/td&gt;
&lt;td&gt;2026-07-29&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;Lost&lt;/strong&gt; — badly&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Rematch, handicap removed&lt;/td&gt;
&lt;td&gt;2026-07-29&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;Lost again&lt;/strong&gt;, by less, for a different reason&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Better-converged measurements&lt;/td&gt;
&lt;td&gt;2026-07-31&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;Lost&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;More honest price list&lt;/td&gt;
&lt;td&gt;2026-08-02&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;Lost&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Most honest price list yet&lt;/td&gt;
&lt;td&gt;2026-08-06&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;Lost by the most&lt;/strong&gt; since the handicaps came off&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Ban the most aggressive setting&lt;/td&gt;
&lt;td&gt;2026-08-06&lt;/td&gt;
&lt;td&gt;Tie on one metric, behind on the other&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Pipeline builds its own winner&lt;/td&gt;
&lt;td&gt;2026-08-09&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;Won&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;Those are the full-loop runs. Several hand-built probes in between are left out, and several of those lost too.&lt;/p&gt;
&lt;p&gt;Read the table as a diagnosis, not a confession — each loss eliminated a suspect. The first proved most of my deficit came from a weighting file the competition used and I didn't, not from my recipe at all. The second proved damage doesn't simply add up: crushing two layers together can hurt far more than crushing each alone, which broke an assumption at the center of my solver. The fifth killed my favorite theory — three rounds of better measurements produced three steps backwards, because the frame I measured in and the frame I shipped in disagreed.&lt;/p&gt;
&lt;p&gt;Each loss was me finding a step nobody had written down. The weighting file and the way damage compounds were both things the preset had been quietly getting right for years without anybody saying so — the pinch of salt your abuelita never mentions because her hand does it automatically. The difference between an accident and a technique is whether somebody writes it down.&lt;/p&gt;
&lt;h3&gt;The recipe I did not expect&lt;/h3&gt;
&lt;p&gt;I assumed measuring would produce something exotic — a wild, uneven allocation no rule of thumb would ever guess. It did, and those were the recipes that lost.&lt;/p&gt;
&lt;p&gt;&lt;img src="https://alberto.codes/vramfit-recipe-shapes.svg" alt="Three bit allocations across 82 layer groups. The preset is a flat 3-bit band. My first attempt is jagged, with 38 groups pushed down to 2-bit, and it lost. The winner is almost flat 3-bit like the preset, with one group at 8-bit, plus a hatched strip marking one tensor kept at higher precision inside 47 of the layers." /&gt;&lt;/p&gt;
&lt;p&gt;Given honest prices, the solver walked back to almost exactly the shape the preset already had. Your abuelita was right about the dish. What the measurement found was the one step she never mentioned — a single tensor inside the layers, kept at higher precision, in 47 places. That strip is the entire difference between what I published and what everyone downloads.&lt;/p&gt;
&lt;p&gt;So I didn't beat the preset by out-thinking it. I beat it by finding the one thing it was missing, and only after five losses told me where not to look. Tracking that down meant working out why the same tensor was fitted ten times worse in my files than in the competition's. I suspected my settings, then my toolchain version. Both innocent. The real cause was upstream and subtle, in how the compression maths behaves when the weighting file has extreme values. Not my bug — but mine to find.&lt;/p&gt;
&lt;h3&gt;Why the losses are public&lt;/h3&gt;
&lt;p&gt;Every one of those losses is dated and numbered, with its receipts. Seventeen data points. Architecture decision records for the choices, an issue trail for the arguments, provenance and evaluation sidecars on the published files.&lt;/p&gt;
&lt;p&gt;I don't think that's normal, and I think it should be. The going standard for a published quantized model is a checksum, a file size, and sometimes a single quality number. A checksum proves the file is the file. It proves nothing about whether the file is any good — and nothing about the five artifacts that didn't make it.&lt;/p&gt;
&lt;p&gt;Weigh the ingredients. Write down the voyage.&lt;/p&gt;
&lt;h2&gt;What I'm still not claiming&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;The winning run was not a cold solve.&lt;/strong&gt; By the time the pipeline built the published artifact, I had hand-discovered three things and handed them to it as inputs: ban the most aggressive setting, protect one tensor across 47 layers, exclude four tensors from the weighting file. The solver reproduced the rest of the layout and made one forced trade of its own. &amp;quot;Measure instead of guessing&amp;quot; is true. &amp;quot;The machine worked it out by itself&amp;quot; is not — the measurements found those rules across five losing attempts, but a human read the measurements.&lt;/p&gt;
&lt;p&gt;The baseline beats me on one metric — how often the top word matches the original — by half a point, and that has never flipped. I let KL carry the ranking anyway, because the top word is a one-bit summary that ignores how wrong the model is when it misses, while KL weights the whole shape. If you always decode greedily, take that half point seriously. If you sample, take the KL.&lt;/p&gt;
&lt;p&gt;And one control was missing when I first published this. When both models got the same four facts wrong, the natural reading is that the mistakes come from the original rather than from either compression. Natural isn't measured. The 93 GB original had never answered those fifteen questions, because it doesn't fit the card. I said so here, &lt;a href="https://github.com/Alberto-Codes/vramfit/issues/143"&gt;tracked it in the open&lt;/a&gt;, and said that if the original got them &lt;em&gt;right&lt;/em&gt;, that would be more interesting than the result I had.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;It got them wrong.&lt;/strong&gt; I ran it the same night. &lt;strong&gt;93 GB&lt;/strong&gt; still doesn't fit &lt;strong&gt;24 GB&lt;/strong&gt;, so most of the model ran on the CPU — about one word every three or four seconds, thirty-five minutes for fifteen answers. It missed the same four facts. It wrote the same function with the same bug, under the same comment promising it had avoided that exact bug. The mistakes come from the original. Neither compression caused them.&lt;/p&gt;
&lt;p&gt;The original scored &lt;strong&gt;20 out of 25&lt;/strong&gt; — one point above both copies. Each compression gave up exactly one further point, and they turned out to be the two points that already separated the copies from each other: one flubbed a size parser, the other expanded an acronym wrong. The original got both right.&lt;/p&gt;
&lt;p&gt;One point is not nothing, and I won't round it to zero. Compression may have cost each copy that point. Fifteen questions, asked once, cannot tell you whether it did — one point out of twenty-five sits inside what the choice of questions alone can move. That's the same weakness that made the 19-19 tie a weak result, and it cuts the same way in both directions. The conversation can't see the difference. The measurement can.&lt;/p&gt;
&lt;p&gt;I'd rather publish the gap named than the story clean. Naming it is also what made it cheap to close.&lt;/p&gt;
&lt;h2&gt;Where it lives&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;The tool:&lt;/strong&gt; &lt;a href="https://github.com/Alberto-Codes/vramfit"&gt;vramfit&lt;/a&gt; — scan, plan, validate, pack. MIT.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;The model:&lt;/strong&gt; &lt;a href="https://huggingface.co/Alberto-Codes/Llama-3_3-Nemotron-Super-49B-v1_5-fit24gib-GGUF"&gt;the packed 49B&lt;/a&gt; — a quantization of &lt;a href="https://huggingface.co/nvidia/Llama-3_3-Nemotron-Super-49B-v1_5"&gt;NVIDIA's Nemotron Super 49B v1_5&lt;/a&gt;, under the NVIDIA Open Model License. Built with Llama.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;The measurements:&lt;/strong&gt; the &lt;a href="https://huggingface.co/datasets/Alberto-Codes/Llama-3_3-Nemotron-Super-49B-v1_5-sensitivity-maps"&gt;sensitivity-map dataset&lt;/a&gt; — the per-layer price list the recipe was solved from.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;The full ledger:&lt;/strong&gt; &lt;a href="https://github.com/Alberto-Codes/vramfit/blob/main/docs/explanation/evaluating-packed-models.md"&gt;all the data points&lt;/a&gt;, every number and every loss.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Coming next: how to fit a model to the card you actually have, and why some layers break and others don't. It's also the second time I've come at quantization from the measurement end — the &lt;a href="https://alberto.codes/blog/2026-03-26-i-ran-turboquant-on-a-vision-model-the-first-output-was-garbage"&gt;first was TurboQuant on a vision model&lt;/a&gt;, where the first output was garbage for a reason no benchmark would have told me.&lt;/p&gt;
</content:encoded>
    </item>
    <item>
      <title>"Two patterns that make production database changes boring"</title>
      <link>https://alberto.codes/blog/2026-06-02-two-patterns-that-make-prod-db-changes-boring</link>
      <guid isPermaLink="true">https://alberto.codes/blog/2026-06-02-two-patterns-that-make-prod-db-changes-boring</guid>
      <pubDate>Tue, 02 Jun 2026 00:00:00 GMT</pubDate>
      <description>"Two patterns that make production database changes safe, auditable, and boring. How SCD Type 2 and maker-checker work — and why config data belongs in a database, not your git repo."</description>
      <content:encoded>&lt;p&gt;Imagine a restaurant where changing the daily special requires a full kitchen renovation. New permits. Contractor review. A week of construction. The fish market calls at 6 AM with beautiful halibut, but the chef can't put it on the board until the renovation crew finishes, the inspector signs off, and the general manager reviews the blueprints.&lt;/p&gt;
&lt;p&gt;That's how some teams treat production configuration changes. A prompt template needs a tweak? Pull request. Code review from an engineer who doesn't evaluate prompts. CI pipeline. Deployment. Change request ticket. The works. By the time the change is live, the moment has passed — like telling the chef the halibut is approved three days after it stopped being fresh.&lt;/p&gt;
&lt;p&gt;The daily special exists because restaurants figured out something important: the kitchen infrastructure and the menu are two different things. You build the kitchen once, to code, with proper ventilation and fire suppression and health inspections. Then the chef changes the specials board whenever the market delivers something worth serving. The kitchen doesn't change. The specials board does.&lt;/p&gt;
&lt;p&gt;Production configuration works the same way. Application logic, database schemas, API contracts — those are the kitchen. Prompt templates, feature flags, business rules, agent skill configs — those are the daily special. And regulated industries figured out how to change the daily special safely decades before most of us wrote our first &lt;code&gt;SELECT&lt;/code&gt; statement.&lt;/p&gt;
&lt;p&gt;Two patterns do most of the work: SCD Type 2 and maker-checker.&lt;/p&gt;
&lt;hr /&gt;
&lt;h2&gt;The two problems people conflate&lt;/h2&gt;
&lt;p&gt;When someone says &amp;quot;put it in code so it goes through the SDLC,&amp;quot; they're conflating two different concerns:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;&lt;strong&gt;Change safety&lt;/strong&gt;: preventing bad changes from reaching production&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Change auditability&lt;/strong&gt;: knowing who changed what, when, and why&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;Code review solves both for &lt;em&gt;code&lt;/em&gt;. But for runtime configuration — the daily specials — it solves neither well. A prompt template buried in a YAML file gets the same review process as a critical algorithm change. The reviewer is checking for merge conflicts and syntax, not whether the prompt actually classifies orders better. And the audit trail is git history, which tells you when the file changed but not which version was active at 3:47 PM on Tuesday when the system started behaving strangely.&lt;/p&gt;
&lt;p&gt;The SDLC is the right process for building the kitchen. It's the wrong process for deciding what goes on the specials board.&lt;/p&gt;
&lt;p&gt;This isn't theoretical:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;In 2008, Jérôme Kerviel cost Société Générale €4.9 billion by accumulating enough system entitlements to approve his own trades — one person, both maker and checker. A database constraint enforcing that the proposer and approver must be different people would have made that physically impossible.&lt;/li&gt;
&lt;li&gt;In 2012, Knight Capital lost $440 million in 45 minutes because a configuration deployment reactivated dormant code and no second person signed off on the change.&lt;/li&gt;
&lt;li&gt;And the classic that every DBA has nightmares about: an UPDATE without a WHERE clause on Black Friday that zeroed out a production table and cost $47M — because the old values were gone the moment the UPDATE committed. With SCD2, there is no UPDATE. Every change is a new row. The previous values are always there.&lt;/li&gt;
&lt;/ul&gt;
&lt;hr /&gt;
&lt;h2&gt;The litmus test: kitchen vs. specials board&lt;/h2&gt;
&lt;p&gt;Here's the question I use: &lt;strong&gt;does this thing change on the same cadence as a code deployment?&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;If yes — it's part of the kitchen. Put it in your repo. Review it. Deploy it.&lt;/p&gt;
&lt;p&gt;If no — it's on the specials board. Put it in a database with SCD2 and maker-checker.&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Thing&lt;/th&gt;
&lt;th&gt;Changes with deploys?&lt;/th&gt;
&lt;th&gt;Verdict&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Application logic&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;Kitchen&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Database schema&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;Kitchen (migrations)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;API contracts&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;Kitchen&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Prompt templates&lt;/td&gt;
&lt;td&gt;No — tuned weekly, sometimes daily&lt;/td&gt;
&lt;td&gt;Specials board&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Feature flags&lt;/td&gt;
&lt;td&gt;No — toggled at runtime&lt;/td&gt;
&lt;td&gt;Specials board&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Business rules / thresholds&lt;/td&gt;
&lt;td&gt;No — adjusted by business users&lt;/td&gt;
&lt;td&gt;Specials board&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Agent skill configs&lt;/td&gt;
&lt;td&gt;No — updated as capabilities evolve&lt;/td&gt;
&lt;td&gt;Specials board&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Notification templates&lt;/td&gt;
&lt;td&gt;No — marketing changes them&lt;/td&gt;
&lt;td&gt;Specials board&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Pricing tiers&lt;/td&gt;
&lt;td&gt;No — product changes them quarterly&lt;/td&gt;
&lt;td&gt;Specials board&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;When a prompt engineer needs to tweak a classifier prompt on Thursday afternoon, they shouldn't need a PR, a code review from an engineer who doesn't understand prompts, a CI pipeline, and a production deployment. They should propose the change, have another prompt engineer approve it, and have the system apply it — with a full audit trail, instant rollback, and zero downtime.&lt;/p&gt;
&lt;p&gt;That's not less rigorous than the SDLC. It's &lt;em&gt;more&lt;/em&gt; rigorous — because the controls are systemic, not procedural. The database constraint enforces dual authorization whether it's Tuesday morning or 2 AM on a holiday. A code review process depends on humans following the rules every time.&lt;/p&gt;
&lt;hr /&gt;
&lt;h2&gt;Pattern 1: SCD Type 2 — the chef's notebook&lt;/h2&gt;
&lt;p&gt;Slowly Changing Dimension Type 2 comes from data warehousing, but the pattern applies anywhere you need a complete history of changes to a record. The idea is simple: you never update or delete. Every change inserts a new row.&lt;/p&gt;
&lt;p&gt;Think of it as the chef's notebook — the one where every daily special ever served is recorded with the date, what was in it, and why it was chosen. The chef doesn't erase yesterday's entry when today's special changes. They write a new line. Six months from now, when someone asks &amp;quot;what were we running the week that got all those compliments,&amp;quot; the answer is right there.&lt;/p&gt;
&lt;pre&gt;&lt;code class="language-sql"&gt;CREATE TABLE prompt_templates (
    id              SERIAL PRIMARY KEY,
    prompt_key      TEXT NOT NULL,
    template_body   TEXT NOT NULL,
    version         INT NOT NULL,

    -- audit columns (inherited from base model in practice)
    created_at      TIMESTAMPTZ NOT NULL DEFAULT now(),
    created_by      TEXT NOT NULL,
    approved_by     TEXT,
    approved_at     TIMESTAMPTZ,
    change_reason   TEXT NOT NULL,

    -- temporal columns
    start_date      TIMESTAMPTZ,            -- set on approval
    end_date        TIMESTAMPTZ,            -- NULL = current
    is_current      BOOLEAN NOT NULL DEFAULT FALSE
);
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;When you change a prompt, you don't &lt;code&gt;UPDATE&lt;/code&gt;. You propose a new row, a second person approves it, and the system expires the old version and activates the new one:&lt;/p&gt;
&lt;pre&gt;&lt;code class="language-sql"&gt;-- Step 1: Propose (created_by fills in, approved_by stays NULL)
INSERT INTO prompt_templates
  (prompt_key, template_body, version, created_by, change_reason)
VALUES
  ('order_classifier', 'Classify the following order...', 3,
   'gordon', 'Reduced false positives on international orders');

-- Step 2: Approve (a different user)
BEGIN;
  UPDATE prompt_templates
     SET end_date = now(), is_current = FALSE
   WHERE prompt_key = 'order_classifier' AND is_current = TRUE;

  UPDATE prompt_templates
     SET approved_by = 'marco', approved_at = now(),
         start_date = now(), is_current = TRUE
   WHERE prompt_key = 'order_classifier' AND version = 3;
COMMIT;
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;What you get for free:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Full audit trail&lt;/strong&gt;: every version that was ever active, who created it, when it was live, and why it changed. Not in a log file somewhere — in the data itself.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Point-in-time queries&lt;/strong&gt;: &amp;quot;What prompt was active at 2026-05-15 14:30 UTC?&amp;quot; is a one-liner: &lt;code&gt;WHERE prompt_key = 'order_classifier' AND start_date &amp;lt;= '2026-05-15 14:30+00' AND (end_date IS NULL OR end_date &amp;gt; '2026-05-15 14:30+00')&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Instant rollback&lt;/strong&gt;: expire the current row and copy any previous version forward. One transaction. No deployment pipeline.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Zero data loss&lt;/strong&gt;: nothing is ever physically deleted. Every state the system has ever been in is recoverable.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Compare this to the git-based approach: to answer &amp;quot;what prompt was active during that incident at 3:47 PM,&amp;quot; you'd need to correlate git commit timestamps with deployment timestamps with the actual moment the running application picked up the new config. With SCD2, the answer is one query.&lt;/p&gt;
&lt;h3&gt;When today's special isn't working: rollback&lt;/h3&gt;
&lt;p&gt;Thursday's halibut special isn't moving. The chef doesn't tear out the stove and reinstall last week's kitchen. They flip back two pages in the notebook and bring back Tuesday's short rib that sold out by 7 PM.&lt;/p&gt;
&lt;p&gt;With SCD2, rollback works the same way. Say version 3 of your prompt is producing bad classifications. You want to go back to version 2 — the one that was working fine last week. Every previous version is still in the table:&lt;/p&gt;
&lt;pre&gt;&lt;code class="language-sql"&gt;BEGIN;
  -- Expire the current version
  UPDATE prompt_templates
     SET end_date = now(), is_current = FALSE
   WHERE prompt_key = 'order_classifier' AND is_current = TRUE;

  -- Bring back version 2 as a new entry
  INSERT INTO prompt_templates
    (prompt_key, template_body, version, created_by, approved_by,
     approved_at, start_date, is_current, change_reason)
  SELECT
    prompt_key, template_body, 4, 'gordon', 'marco',
    now(), now(), TRUE,
    'Rollback to v2: v3 increased false positives on international orders'
  FROM prompt_templates
  WHERE prompt_key = 'order_classifier' AND version = 2;
COMMIT;
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Version 3 isn't deleted — it's preserved with timestamps showing exactly when it was active and when it was rolled back. Version 4 is a copy of version 2's content with a new &lt;code&gt;change_reason&lt;/code&gt; explaining &lt;em&gt;why&lt;/em&gt;. If an auditor asks what happened, you have the full story: v3 went live at 14:00, caused issues, rolled back at 15:22, and here's what v3 was trying to do.&lt;/p&gt;
&lt;p&gt;No deployment queue. No waiting for CI. No merge conflicts because someone else pushed to main in the meantime. The application picks up the rolled-back config on its next read — the way a restaurant picks up the new specials board without rebuilding the dining room.&lt;/p&gt;
&lt;hr /&gt;
&lt;h2&gt;Pattern 2: Maker-checker — the sous chef and the head chef&lt;/h2&gt;
&lt;p&gt;Also called the four-eyes principle, maker-checker splits every sensitive operation into two steps: one person proposes a change, a different person approves it. The system physically prevents single-actor completion regardless of permission level.&lt;/p&gt;
&lt;p&gt;Every kitchen runs on a version of this. The sous chef writes the special on a ticket. The head chef tastes the dish, approves it, and &lt;em&gt;then&lt;/em&gt; it goes on the board. Two sets of hands, not because the sous chef can't cook — because the kitchen is designed so nothing reaches the dining room without a second check.&lt;/p&gt;
&lt;p&gt;&lt;img src="https://alberto.codes/maker-checker-workflow.svg" alt="Maker-checker workflow: Gordon writes the special, Marco Pierre White tastes and approves or rejects — a different person must review" /&gt;&lt;/p&gt;
&lt;p&gt;You don't need a separate approvals table. The maker-checker columns live on the same row as the data — the same &lt;code&gt;prompt_templates&lt;/code&gt; table from above. A row with &lt;code&gt;approved_by IS NULL&lt;/code&gt; is a pending proposal. A row with &lt;code&gt;approved_by&lt;/code&gt; filled in and &lt;code&gt;start_date&lt;/code&gt; set is live. The database enforces the constraint:&lt;/p&gt;
&lt;pre&gt;&lt;code class="language-sql"&gt;CONSTRAINT different_actors
    CHECK (created_by != approved_by OR approved_by IS NULL)
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;That's the whole point. The database itself enforces that the maker cannot be the checker. This isn't a policy in a wiki that people follow when they remember — it's a constraint that the system cannot violate.&lt;/p&gt;
&lt;p&gt;The workflow maps to the columns directly:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;&lt;strong&gt;Propose&lt;/strong&gt;: &lt;code&gt;created_by&lt;/code&gt; and &lt;code&gt;created_at&lt;/code&gt; fill in. &lt;code&gt;approved_by&lt;/code&gt; stays NULL. The row exists but isn't active.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Review&lt;/strong&gt;: a different user approves — &lt;code&gt;approved_by&lt;/code&gt; and &lt;code&gt;approved_at&lt;/code&gt; fill in.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Activate&lt;/strong&gt;: &lt;code&gt;start_date&lt;/code&gt; is set, &lt;code&gt;is_current&lt;/code&gt; flips to TRUE, and the previous version gets an &lt;code&gt;end_date&lt;/code&gt;.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;This is how most banks process wire transfers. How most hospitals manage medication orders. The pattern is so well-established that regulators don't just recommend it — they mandate it.&lt;/p&gt;
&lt;hr /&gt;
&lt;h2&gt;What the regulators actually require&lt;/h2&gt;
&lt;p&gt;If the instinct is &amp;quot;but compliance won't allow database changes in prod,&amp;quot; it's worth checking what the compliance frameworks actually say — because these patterns are exactly what they describe.&lt;/p&gt;
&lt;p&gt;Health codes don't require a renovation to change the menu — they require that the kitchen is built to code, and the daily decisions flow through it safely.&lt;/p&gt;
&lt;p&gt;Regulatory frameworks work the same way:&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Standard&lt;/th&gt;
&lt;th&gt;What it requires&lt;/th&gt;
&lt;th&gt;Retention&lt;/th&gt;
&lt;th&gt;What SCD2 + maker-checker gives you&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;SOX §302/§404&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Internal controls, audit trails for financial systems&lt;/td&gt;
&lt;td&gt;7 years&lt;/td&gt;
&lt;td&gt;Immutable history + segregation of duties&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;HIPAA (21 CFR Part 11)&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Audit trails: who modified data, when, what changed&lt;/td&gt;
&lt;td&gt;6 years&lt;/td&gt;
&lt;td&gt;&lt;code&gt;created_by&lt;/code&gt;, &lt;code&gt;approved_by&lt;/code&gt;, &lt;code&gt;start_date&lt;/code&gt;, &lt;code&gt;change_reason&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;PCI-DSS v4.0&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Logging of all changes to system components&lt;/td&gt;
&lt;td&gt;12 months&lt;/td&gt;
&lt;td&gt;Approval workflow + full version history&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;NIST 800-53 AU&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Event logging: who, what, when, where, outcome&lt;/td&gt;
&lt;td&gt;Policy-defined&lt;/td&gt;
&lt;td&gt;Append-only records, no physical deletes&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;EU DORA&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Change control with full audit trails for ICT systems&lt;/td&gt;
&lt;td&gt;Per-contract&lt;/td&gt;
&lt;td&gt;Maker-checker workflow + SCD2 trail&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;GDPR Article 30&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Records of processing activities: what and why&lt;/td&gt;
&lt;td&gt;Duration of processing&lt;/td&gt;
&lt;td&gt;&lt;code&gt;change_reason&lt;/code&gt; field + complete history&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;The pattern across all of them: &lt;em&gt;who&lt;/em&gt; changed &lt;em&gt;what&lt;/em&gt;, &lt;em&gt;when&lt;/em&gt;, and &lt;em&gt;why&lt;/em&gt; — with immutable records and independent verification. That's SCD2 + maker-checker by construction.&lt;/p&gt;
&lt;hr /&gt;
&lt;h2&gt;Putting it together&lt;/h2&gt;
&lt;p&gt;The chef's notebook after a week of specials:&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;v&lt;/th&gt;
&lt;th&gt;special&lt;/th&gt;
&lt;th&gt;created&lt;/th&gt;
&lt;th&gt;approved&lt;/th&gt;
&lt;th&gt;start&lt;/th&gt;
&lt;th&gt;end&lt;/th&gt;
&lt;th&gt;cur&lt;/th&gt;
&lt;th&gt;reason&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;Pan-seared halibut, lemon beurre blanc&lt;/td&gt;
&lt;td&gt;Gordon&lt;/td&gt;
&lt;td&gt;Marco&lt;/td&gt;
&lt;td&gt;Mon&lt;/td&gt;
&lt;td&gt;Wed&lt;/td&gt;
&lt;td&gt;no&lt;/td&gt;
&lt;td&gt;Fresh halibut from market&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;Wagyu beef tartare, truffle vinaigrette&lt;/td&gt;
&lt;td&gt;Gordon&lt;/td&gt;
&lt;td&gt;Marco&lt;/td&gt;
&lt;td&gt;Wed&lt;/td&gt;
&lt;td&gt;Fri&lt;/td&gt;
&lt;td&gt;no&lt;/td&gt;
&lt;td&gt;Wagyu delivery came in&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;td&gt;Lobster risotto, saffron cream&lt;/td&gt;
&lt;td&gt;Marco&lt;/td&gt;
&lt;td&gt;Gordon&lt;/td&gt;
&lt;td&gt;Fri&lt;/td&gt;
&lt;td&gt;Sat&lt;/td&gt;
&lt;td&gt;no&lt;/td&gt;
&lt;td&gt;Weekend prix fixe&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;4&lt;/td&gt;
&lt;td&gt;Wagyu beef tartare, truffle vinaigrette&lt;/td&gt;
&lt;td&gt;Gordon&lt;/td&gt;
&lt;td&gt;Marco&lt;/td&gt;
&lt;td&gt;Sat&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;yes&lt;/td&gt;
&lt;td&gt;Rollback: risotto undersold&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;Four rows. Every special that ran, who wrote it, who approved it, when it was live, and why it changed. Saturday's special is Wednesday's wagyu brought back — the risotto didn't sell, so they flipped the notebook back two pages. Nothing erased. The health inspector can read it top to bottom and see the whole week.&lt;/p&gt;
&lt;p&gt;The specials board changes. The kitchen stays.&lt;/p&gt;
&lt;hr /&gt;
&lt;h2&gt;What's next: making the SQL disappear&lt;/h2&gt;
&lt;p&gt;The raw SQL in this post is worth understanding — you should know what's happening at the database level. But in a real application, nobody should be writing expire-and-insert transactions by hand.&lt;/p&gt;
&lt;p&gt;In a follow-up, I'll show how to wrap these patterns in a domain layer using SQLModel, the Unit of Work pattern, and ports and adapters. The idea: a base model carries all the audit and temporal columns — &lt;code&gt;created_at&lt;/code&gt;, &lt;code&gt;created_by&lt;/code&gt;, &lt;code&gt;approved_by&lt;/code&gt;, &lt;code&gt;start_date&lt;/code&gt;, &lt;code&gt;end_date&lt;/code&gt; — and your domain models inherit from it. &lt;code&gt;PromptTemplate&lt;/code&gt; just adds &lt;code&gt;prompt_key&lt;/code&gt; and &lt;code&gt;template_body&lt;/code&gt;. The SCD2 versioning, maker-checker enforcement, and approval workflow happen automatically in the infrastructure layer. Your application code never writes a temporal query directly — the same way a well-run kitchen handles food safety without the chef thinking about health codes on every plate.&lt;/p&gt;
&lt;hr /&gt;
&lt;h2&gt;Key takeaways&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;SCD Type 2&lt;/strong&gt; gives you a complete, immutable history of every configuration change — who, what, when, why — queryable by timestamp.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Maker-checker&lt;/strong&gt; enforces dual authorization at the database level, not as a policy people can skip.&lt;/li&gt;
&lt;li&gt;Together, these patterns map directly to the audit trail requirements of SOX, HIPAA, PCI-DSS, NIST 800-53, DORA, and GDPR — giving you the building blocks by construction, not by procedure.&lt;/li&gt;
&lt;li&gt;The litmus test: if it changes on a different cadence than your code deployments, it's data, not code. Treat it accordingly.&lt;/li&gt;
&lt;li&gt;Rollback is one transaction, not a revert-rebuild-redeploy cycle.&lt;/li&gt;
&lt;li&gt;The fear of production isn't solved by avoiding production. It's solved by building the kitchen to code so the daily specials can change safely.&lt;/li&gt;
&lt;/ul&gt;
&lt;hr /&gt;
&lt;h2&gt;Further reading&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href="https://en.wikipedia.org/wiki/2008_Soci%C3%A9t%C3%A9_G%C3%A9n%C3%A9rale_trading_loss"&gt;Société Générale: the €4.9B rogue trader&lt;/a&gt; — one person approved his own trades; dual authorization would have stopped it&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.henricodolfing.ch/en/case-study-4-the-440-million-software-error-at-knight-capital/"&gt;The Knight Capital disaster&lt;/a&gt; — $440M lost in 45 minutes from an unreviewed config deployment&lt;/li&gt;
&lt;li&gt;&lt;a href="https://thecodeforge.io/database/sql-insert-update-delete/"&gt;The $47M UPDATE without a WHERE clause&lt;/a&gt; — destructive UPDATE zeroed out production data; SCD2's append-only model prevents this by design&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.analyticsengineering.com/resources/slowly-changing-dimensions-type-2-explained"&gt;SCD Type 2 explained&lt;/a&gt; — comprehensive walkthrough of the pattern&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.artie.com/blogs/how-to-build-audit-logs-using-cdc-and-scd-type-2"&gt;Building audit logs with CDC and SCD Type 2&lt;/a&gt; — combining change data capture with SCD2&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.opcito.com/blogs/maker-checker-implementation-guide-for-secure-fintech-systems"&gt;Maker-checker implementation guide for fintech&lt;/a&gt; — real-world banking implementation&lt;/li&gt;
&lt;li&gt;&lt;a href="https://csf.tools/reference/nist-sp-800-53/r5/au/"&gt;NIST 800-53 AU controls&lt;/a&gt; — the audit and accountability family&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.liquibase.com/blog/what-is-the-digital-operational-resilience-act-dora-what-financial-database-developer-teams-need-to-know"&gt;DORA and database teams&lt;/a&gt; — what DORA means for database change management&lt;/li&gt;
&lt;li&gt;&lt;a href="https://launchdarkly.com/blog/feature-flags-vs-deployment-automation-vs-config-files/"&gt;Feature flags vs. config files vs. deployment&lt;/a&gt; — when each approach is appropriate&lt;/li&gt;
&lt;li&gt;&lt;a href="https://martinfowler.com/articles/feature-toggles.html"&gt;Martin Fowler on feature toggles&lt;/a&gt; — the canonical reference on runtime configuration&lt;/li&gt;
&lt;/ul&gt;
</content:encoded>
    </item>
    <item>
      <title>"From one model to seven: Making TurboQuant model-portable"</title>
      <link>https://alberto.codes/blog/2026-03-31-from-one-model-to-seven-making-turboquant-model-portable</link>
      <guid isPermaLink="true">https://alberto.codes/blog/2026-03-31-from-one-model-to-seven-making-turboquant-model-portable</guid>
      <pubDate>Tue, 31 Mar 2026 00:00:00 GMT</pubDate>
      <description>"turboquant-vllm started as a Molmo2-only proof of concept. v1.3.0 validates seven model families — but getting there meant rewriting Triton kernels for non-standard head dimensions and teaching the cache about sliding window attention."</description>
      <content:encoded>&lt;p&gt;A compression algorithm that only works on one model isn't a tool — it's a demo. When &lt;a href="https://alberto.codes/blog/2026-03-27-paper-to-pypi-in-72-hours-building-the-first-turboquant-vllm-plugin"&gt;turboquant-vllm v1.0.0&lt;/a&gt; shipped, it was validated on exactly one architecture: Molmo2. The algorithm worked, the numbers were real (3.76x KV compression, ~97% cosine similarity), but every model has its own attention geometry. Head dimensions vary. Some layers use sliding windows. Triton kernels crash when you hand them a dimension that isn't a power of two.&lt;/p&gt;
&lt;p&gt;v1.3.0 validates seven model families: Molmo2, Llama 3.1, Mistral 7B, Qwen2.5, Phi-3-mini, Phi-4, Gemma-2, and Gemma-3. Getting there required two things: fused kernels that actually perform well in production (v1.2.0), and kernel-level changes to handle the architectural diversity across those families (v1.3.0).&lt;/p&gt;
&lt;h2&gt;Fused kernels: decompress and attend in one pass&lt;/h2&gt;
&lt;p&gt;The v1.1.0 architecture had a clean separation: decompress the KV cache from TQ4 to FP16 in HBM, then run standard attention on the decompressed data. Clean, but wasteful — every decode step wrote decompressed values to HBM just to read them back immediately.&lt;/p&gt;
&lt;p&gt;v1.2.0 introduced fused paged TQ4 kernels that eliminate that round trip. The decode kernel reads compressed blocks directly from vLLM's page table, decompresses in SRAM (nibble unpack → centroid gather → norm scale), and computes Q@K^T with online softmax — all in a single Triton kernel. No HBM writes of decompressed cache.&lt;/p&gt;
&lt;p&gt;&lt;img src="https://alberto.codes/turboquant-fused-kernel-flow.svg" alt="Fused kernel data flow: v1.1 separate passes vs v1.2 single fused pass" /&gt;&lt;/p&gt;
&lt;p&gt;The concrete impact: HBM traffic drops from 1,160 to 136 bytes per token — an 8.5x reduction. A separate INT8 tensor core prefill kernel handles the prefill path with the same fusion strategy.&lt;/p&gt;
&lt;p&gt;This wasn't just a performance optimization. It was the foundation for everything that followed. CUDA graph buffer pre-allocation eliminated kernel launch latency for decode. Feature gating allowed the fused path to coexist with the original decompress-first path as a fallback. And when container benchmarks (Experiments 022–023) surfaced OOM bugs in the scratch buffers, the fixes (v1.2.1, v1.2.2) landed in hours because the architecture was clean enough to patch confidently.&lt;/p&gt;
&lt;h2&gt;The portability problem: Triton and power-of-two&lt;/h2&gt;
&lt;p&gt;Here's what I didn't appreciate until I tried running TQ4 on Phi-3-mini: Triton's &lt;code&gt;tl.arange&lt;/code&gt; requires power-of-two ranges. Molmo2 has head_dim=128 — a power of two. Phi-3-mini has head_dim=96. The kernels crashed at compile time.&lt;/p&gt;
&lt;p&gt;The fix sounds simple — pad to the next power of two and mask the boundary. In practice, it touched all five Triton kernels: &lt;code&gt;flash_attention&lt;/code&gt;, &lt;code&gt;flash_attention_tq4&lt;/code&gt;, &lt;code&gt;flash_attention_tq4_kv&lt;/code&gt;, &lt;code&gt;tq4_compress&lt;/code&gt;, and &lt;code&gt;tq4_decompress&lt;/code&gt;. Each needed a &lt;code&gt;_next_pow2&lt;/code&gt; helper, &lt;code&gt;HEAD_DIM_PAD&lt;/code&gt;/&lt;code&gt;HALF_D_PAD&lt;/code&gt; constants, and &lt;code&gt;d_mask&lt;/code&gt; boundary guards to prevent out-of-bounds reads during the fused decompression step.&lt;/p&gt;
&lt;p&gt;Gemma-2 and Gemma-3 added another dimension — literally. head_dim=256 required tuning the flash attention autotune search space (adding &lt;code&gt;BLOCK_M=32&lt;/code&gt; for SRAM optimization). The kernel works, but the autotune cost was real.&lt;/p&gt;
&lt;p&gt;The throughput penalty for non-pow2 dimensions is ~5–15%, which is honest and documented. For head_dim=128 models — the majority — there's zero penalty.&lt;/p&gt;
&lt;h2&gt;Sliding window attention: not all layers compress equally&lt;/h2&gt;
&lt;p&gt;Gemma models use mixed attention: some layers are global (full context), others use sliding windows (fixed context window with cache eviction). Compressing a sliding window layer's cache breaks the eviction semantics — the cache needs to discard old entries, not keep them in compressed form.&lt;/p&gt;
&lt;p&gt;The solution is a bypass. When &lt;code&gt;CompressedDynamicCache&lt;/code&gt; encounters a layer with &lt;code&gt;is_sliding=True&lt;/code&gt; (from HuggingFace's &lt;code&gt;DynamicSlidingWindowLayer&lt;/code&gt;), it skips compression entirely and delegates to the original cache update. Global layers compress normally.&lt;/p&gt;
&lt;p&gt;&lt;img src="https://alberto.codes/turboquant-model-portability.svg" alt="Model portability: how head_dim and attention type determine the kernel path" /&gt;&lt;/p&gt;
&lt;p&gt;The implementation required &lt;code&gt;None&lt;/code&gt; padding in the compressed key/value lists for SWA gaps (to keep layer indices aligned), SWA-aware guards in &lt;code&gt;get_compressed&lt;/code&gt;, &lt;code&gt;get_seq_length&lt;/code&gt;, and &lt;code&gt;compression_stats&lt;/code&gt;, and a diagnostic warning when a Gemma-family config is detected but the cache was created without the model config (which means no SWA metadata). Thirteen new tests cover the bypass, warnings, and downstream guards.&lt;/p&gt;
&lt;h2&gt;Verification at scale&lt;/h2&gt;
&lt;p&gt;Model-by-model validation needed a repeatable process, not ad hoc benchmarking. The &lt;code&gt;verify&lt;/code&gt; CLI (&lt;code&gt;python -m turboquant_vllm.verify --model &amp;lt;name&amp;gt; --bits 4&lt;/code&gt;) loads any HuggingFace model, runs TQ4 compression on random Gaussian input, and reports per-layer cosine similarity against the uncompressed output. Pass threshold: 0.99.&lt;/p&gt;
&lt;p&gt;All eight regression models pass:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Molmo2-4B&lt;/strong&gt; — VLM, head_dim=128, cosine 0.9951&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Llama 3.1 8B&lt;/strong&gt; — text, head_dim=128, lossless at temperature=0&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Mistral 7B&lt;/strong&gt; — text, head_dim=128, lossless at temperature=0&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Qwen2.5-3B&lt;/strong&gt; — text, head_dim=128, cosine ≥0.99&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Phi-3-mini&lt;/strong&gt; — text, head_dim=96 (non-pow2), cosine 0.9952&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Phi-4&lt;/strong&gt; — text, head_dim=128, cosine ≥0.99&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Gemma-2-2b&lt;/strong&gt; — text, head_dim=256 + SWA, cosine 0.9951&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Gemma-3-4B-it&lt;/strong&gt; — text, head_dim=256 + SWA, cosine 0.9951&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Experiment 024 pushed further: Llama 3.1 and Mistral 7B through the full vLLM backend with zero code changes. Short prompts, 960-token passage comprehension, 5-turn conversations with 1,200+ KV tokens — all 6/6 PASS on both models. KV capacity: 1.88x advantage over FP8 baseline. At 16K context, TQ4 serves 6x concurrent requests versus baseline's 3x.&lt;/p&gt;
&lt;h2&gt;What this means&lt;/h2&gt;
&lt;p&gt;turboquant-vllm is no longer a single-model proof of concept. If your model uses head_dim 64–256 and runs on vLLM, there's a reasonable chance TQ4 compression works out of the box. The verify CLI takes thirty seconds to check.&lt;/p&gt;
&lt;p&gt;The fused kernels, the non-pow2 padding, the SWA bypass — these are the kind of changes that don't show up in a changelog but determine whether a tool survives contact with the real model ecosystem. v1.0.0 proved the algorithm. v1.3.0 proved the engineering.&lt;/p&gt;
&lt;h2&gt;What's next&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Upstream vLLM contribution&lt;/strong&gt; — the fused paged kernel architecture is the candidate for contribution.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Flash Attention kernel fusion&lt;/strong&gt; — full multi-layer correctness for the fused path, reducing decode overhead further.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;VL-Cache stacking&lt;/strong&gt; — combining TQ4 KV compression with token pruning for multiplicative savings on VLMs.&lt;/li&gt;
&lt;/ul&gt;
&lt;blockquote&gt;
&lt;p&gt;Note 2026-09-06: upstream has since shipped TurboQuant in &lt;a href="https://github.com/vllm-project/vllm/pull/38479"&gt;vllm-project/vllm#38479&lt;/a&gt;, and the feature request, &lt;a href="https://github.com/vllm-project/vllm/issues/38201"&gt;vllm-project/vllm#38201&lt;/a&gt;, is closed.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;hr /&gt;
&lt;p&gt;&lt;em&gt;&lt;a href="https://pypi.org/project/turboquant-vllm/"&gt;PyPI&lt;/a&gt; | &lt;a href="https://alberto-codes.github.io/turboquant-vllm/"&gt;Docs&lt;/a&gt; | &lt;a href="https://github.com/Alberto-Codes/turboquant-vllm"&gt;GitHub&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;
</content:encoded>
    </item>
    <item>
      <title>Serve compressed VLM inference from a container</title>
      <link>https://alberto.codes/blog/2026-03-28-serve-compressed-vlm-inference-from-a-container</link>
      <guid isPermaLink="true">https://alberto.codes/blog/2026-03-28-serve-compressed-vlm-inference-from-a-container</guid>
      <pubDate>Sat, 28 Mar 2026 00:00:00 GMT</pubDate>
      <description>"Build a container image with turboquant-vllm baked in, serve a vision-language model with 3.76x KV cache compression, and verify it works — in under five minutes."</description>
      <content:encoded>&lt;p&gt;The &lt;a href="https://alberto.codes/blog/2026-03-27-paper-to-pypi-in-72-hours-building-the-first-turboquant-vllm-plugin"&gt;first turboquant-vllm release&lt;/a&gt; proved the algorithm works — &lt;code&gt;pip install&lt;/code&gt;, one flag, 3.76x KV cache compression. But if you've ever set up a GPU inference environment from scratch, you know the real friction isn't the model or the framework. It's the CUDA toolkit version, the driver compatibility matrix, the pip packages that refuse to coexist.&lt;/p&gt;
&lt;p&gt;v1.1.0 ships a &lt;code&gt;Containerfile&lt;/code&gt; that eliminates that entire setup. Build the image once, and every run starts from a known-good state — vLLM, CUDA runtime, and the TQ4 compression plugin verified at build time.&lt;/p&gt;
&lt;p&gt;This guide walks through building the container, serving a vision-language model with compressed inference, verifying it works, and optionally running it as a persistent systemd service.&lt;/p&gt;
&lt;p&gt;&lt;img src="https://alberto.codes/turboquant-container-architecture.svg" alt="turboquant-vllm container architecture" /&gt;&lt;/p&gt;
&lt;h2&gt;Prerequisites&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;An NVIDIA GPU&lt;/strong&gt; with drivers installed (tested on RTX 4090, 24 GB). AMD ROCm also works — adjust the device flag.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Podman or Docker.&lt;/strong&gt; Commands below use Podman. For Docker, swap &lt;code&gt;podman&lt;/code&gt; for &lt;code&gt;docker&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Enough VRAM.&lt;/strong&gt; Molmo2-8B needs ~24 GB at 6K context with &lt;code&gt;--gpu-memory-utilization 0.90&lt;/code&gt;. Molmo2-4B fits with longer contexts on the same card.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2&gt;Build the image&lt;/h2&gt;
&lt;p&gt;Clone the repo and build:&lt;/p&gt;
&lt;pre&gt;&lt;code class="language-bash"&gt;git clone https://github.com/Alberto-Codes/turboquant-vllm.git
cd turboquant-vllm
podman build -t vllm-turboquant -f infra/Containerfile.vllm .
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;The &lt;code&gt;Containerfile&lt;/code&gt; does two things: installs &lt;code&gt;turboquant-vllm&lt;/code&gt; from PyPI into the official vLLM image, then verifies the plugin entry point registered correctly. If the entry point check fails, the build fails — you won't discover a misconfigured plugin at runtime.&lt;/p&gt;
&lt;pre&gt;&lt;code class="language-dockerfile"&gt;FROM docker.io/vllm/vllm-openai:v0.18.0
ARG TURBOQUANT_VERSION=1.1.0

RUN pip install --no-cache-dir &amp;quot;turboquant-vllm[vllm]==${TURBOQUANT_VERSION}&amp;quot;

RUN python3 -c &amp;quot;\
import importlib.metadata; \
eps = [e for e in importlib.metadata.entry_points(group='vllm.general_plugins') \
       if e.name == 'tq4_backend']; \
assert len(eps) == 1, 'TQ4 entry point not found'; \
print(f'turboquant-vllm {importlib.metadata.version(\&amp;quot;turboquant-vllm\&amp;quot;)} — plugin verified')&amp;quot;
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;The &lt;code&gt;TURBOQUANT_VERSION&lt;/code&gt; build arg defaults to &lt;code&gt;1.1.0&lt;/code&gt;. Override it for future versions without touching the file.&lt;/p&gt;
&lt;h2&gt;Start the server&lt;/h2&gt;
&lt;pre&gt;&lt;code class="language-bash"&gt;podman run --rm \
  --device nvidia.com/gpu=all \
  --shm-size=8g \
  -v vllm-models:/root/.cache/huggingface \
  -p 8000:8000 \
  vllm-turboquant \
  --model allenai/Molmo2-8B \
  --attention-backend CUSTOM \
  --dtype auto \
  --max-model-len 6144 \
  --max-num-batched-tokens 6144 \
  --enforce-eager \
  --gpu-memory-utilization 0.90
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;One flag does all the work: &lt;code&gt;--attention-backend CUSTOM&lt;/code&gt;. This tells vLLM to use the TQ4 backend instead of its default attention implementation. Everything else — model loading, tokenization, the OpenAI-compatible API — stays exactly the same.&lt;/p&gt;
&lt;p&gt;The named volume (&lt;code&gt;vllm-models&lt;/code&gt;) caches model weights between container restarts. Multi-gigabyte checkpoints download once.&lt;/p&gt;
&lt;h2&gt;Verify compression is active&lt;/h2&gt;
&lt;p&gt;Watch the container logs for the backend confirmation:&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;INFO [cuda.py:257] Using AttentionBackendEnum.CUSTOM backend.
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;If you see &lt;code&gt;FLASH_ATTN&lt;/code&gt; or &lt;code&gt;XFORMERS&lt;/code&gt; instead, the plugin didn't register. Rebuild the image and check the entry point verification passed.&lt;/p&gt;
&lt;p&gt;You can also confirm from inside a running container:&lt;/p&gt;
&lt;pre&gt;&lt;code class="language-bash"&gt;podman exec &amp;lt;container-id&amp;gt; python3 -c &amp;quot;
from turboquant_vllm.vllm import TQ4AttentionBackend
import importlib.metadata
v = importlib.metadata.version('turboquant-vllm')
print(f'turboquant-vllm {v} — plugin loaded')
&amp;quot;
&lt;/code&gt;&lt;/pre&gt;
&lt;h2&gt;Query the API&lt;/h2&gt;
&lt;p&gt;The container exposes the standard vLLM OpenAI-compatible API. Nothing changes on the client side:&lt;/p&gt;
&lt;pre&gt;&lt;code class="language-bash"&gt;curl -s http://localhost:8000/v1/chat/completions \
  -H &amp;quot;Content-Type: application/json&amp;quot; \
  -d '{
    &amp;quot;model&amp;quot;: &amp;quot;allenai/Molmo2-8B&amp;quot;,
    &amp;quot;messages&amp;quot;: [{&amp;quot;role&amp;quot;: &amp;quot;user&amp;quot;, &amp;quot;content&amp;quot;: &amp;quot;Describe this scene&amp;quot;}],
    &amp;quot;max_tokens&amp;quot;: 256
  }' | python3 -m json.tool
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Clients don't know — and don't need to know — that the KV cache is 3.76x compressed behind the API.&lt;/p&gt;
&lt;h2&gt;Persistent deployment with Quadlet&lt;/h2&gt;
&lt;p&gt;For production, Quadlet manages the container as a systemd service. Create &lt;code&gt;~/.config/containers/systemd/vllm-turboquant.container&lt;/code&gt;:&lt;/p&gt;
&lt;pre&gt;&lt;code class="language-ini"&gt;[Container]
Image=localhost/vllm-turboquant:latest
ContainerName=vllm-tq
SecurityLabelDisable=true
ShmSize=8g
AddDevice=nvidia.com/gpu=all
Exec=allenai/Molmo2-8B \
    --attention-backend CUSTOM \
    --dtype auto \
    --max-model-len 6144 \
    --max-num-batched-tokens 6144 \
    --enforce-eager \
    --gpu-memory-utilization 0.90
Volume=vllm-models.volume:/root/.cache/huggingface
PublishPort=8000:8000
HealthCmd=bash -c 'echo &amp;gt; /dev/tcp/localhost/8000'
HealthInterval=30s
HealthTimeout=10s
HealthRetries=5
HealthStartPeriod=300s

[Service]
Restart=always
TimeoutStartSec=900

[Install]
WantedBy=default.target
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Reload and start:&lt;/p&gt;
&lt;pre&gt;&lt;code class="language-bash"&gt;systemctl --user daemon-reload
systemctl --user start vllm-turboquant
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;The health check gives the model up to five minutes to load weights before marking the service unhealthy. &lt;code&gt;Restart=always&lt;/code&gt; handles crashes and GPU driver hiccups automatically.&lt;/p&gt;
&lt;h2&gt;Troubleshooting&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Build fails at entry point verification.&lt;/strong&gt; The vLLM base image version may not match turboquant-vllm's requirements. Check &lt;a href="https://pypi.org/project/turboquant-vllm/"&gt;PyPI&lt;/a&gt; for supported vLLM versions.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Container starts but uses FLASH_ATTN instead of CUSTOM.&lt;/strong&gt; Confirm &lt;code&gt;--attention-backend CUSTOM&lt;/code&gt; is in your run command. In Quadlet, it goes in the &lt;code&gt;Exec=&lt;/code&gt; line after the model name.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;OOM during prefill.&lt;/strong&gt; TurboQuant compresses the KV cache, not model weights or activations. Peak memory during prefill is activation-dominated — compression savings appear during generation. Lower &lt;code&gt;--max-model-len&lt;/code&gt; or use a smaller model variant.&lt;/p&gt;
&lt;h2&gt;What you learned&lt;/h2&gt;
&lt;p&gt;You built a container image with turboquant-vllm baked in, served a vision-language model with 3.76x KV cache compression, verified the plugin was active, and optionally deployed it as a persistent systemd service — all from one &lt;code&gt;Containerfile&lt;/code&gt; and one CLI flag.&lt;/p&gt;
&lt;p&gt;The full API reference and additional usage guides are on the &lt;a href="https://alberto-codes.github.io/turboquant-vllm/"&gt;documentation site&lt;/a&gt;.&lt;/p&gt;
</content:encoded>
    </item>
    <item>
      <title>"Paper to PyPI in 72 hours: Building the first TurboQuant vLLM plugin"</title>
      <link>https://alberto.codes/blog/2026-03-27-paper-to-pypi-in-72-hours-building-the-first-turboquant-vllm-plugin</link>
      <guid isPermaLink="true">https://alberto.codes/blog/2026-03-27-paper-to-pypi-in-72-hours-building-the-first-turboquant-vllm-plugin</guid>
      <pubDate>Fri, 27 Mar 2026 00:00:00 GMT</pubDate>
      <description>"Google published TurboQuant at ICLR 2026 for text models. 72 hours later, turboquant-vllm was on PyPI — the first implementation validated on vision-language models and the first vLLM plugin. One flag to enable, 3.76x KV cache compression."</description>
      <content:encoded>&lt;p&gt;There's a difference between nailing a recipe at home and running it on a restaurant line. At home you control the heat, the timing, the single plate going out. On the line, you need it to work with different stoves, multiple tickets firing at once, and a kitchen that wasn't built around your dish. The &lt;a href="https://alberto.codes/blog/2026-03-26-i-ran-turboquant-on-a-vision-model-the-first-output-was-garbage"&gt;first post&lt;/a&gt; was the home kitchen version — implementing TurboQuant from the paper, finding what works and what breaks. This post is about getting it on the line.&lt;/p&gt;
&lt;p&gt;Google published the &lt;a href="https://arxiv.org/abs/2504.19874"&gt;TurboQuant paper&lt;/a&gt; on March 24, 2026. By March 27, &lt;code&gt;turboquant-vllm&lt;/code&gt; was on PyPI serving compressed video inference through vLLM's OpenAI-compatible API. One flag to enable:&lt;/p&gt;
&lt;pre&gt;&lt;code class="language-bash"&gt;pip install turboquant-vllm[vllm]
vllm serve allenai/Molmo2-8B --attention-backend CUSTOM
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;3.76x KV cache compression. Near-identical output quality. No code changes.&lt;/p&gt;
&lt;p&gt;This post is about the production journey — the decisions that turned a &lt;a href="https://alberto.codes/blog/2026-03-26-i-ran-turboquant-on-a-vision-model-the-first-output-was-garbage"&gt;research implementation&lt;/a&gt; into a pip-installable plugin in 72 hours, and why nobody else has tested TurboQuant on vision-language models.&lt;/p&gt;
&lt;h2&gt;The gap nobody filled&lt;/h2&gt;
&lt;p&gt;Every other TurboQuant implementation I could find — and there are &lt;a href="https://github.com/search?q=turboquant&amp;amp;type=repositories"&gt;several&lt;/a&gt; — tests exclusively on text models: Qwen, Gemma, Mistral, Llama. Google's own paper benchmarks on Gemma, Mistral, and Llama-3.1-8B. Text only.&lt;/p&gt;
&lt;p&gt;Vision-language models are a harder test case. A 12-second video clip through Molmo2-4B produces ~11,000 visual tokens — 10x longer than typical text prompts. That means 10x more KV cache memory, 10x more opportunities for precision bugs to compound across 36 transformer layers.&lt;/p&gt;
&lt;p&gt;The existing VLM KV cache compression literature takes an entirely different approach: token pruning and sparsification (VL-Cache, Dynamic-LLaVA, ZipVL). These methods decide which tokens to &lt;em&gt;discard&lt;/em&gt;. TurboQuant compresses the tokens you &lt;em&gt;keep&lt;/em&gt;. They're complementary — you could stack TurboQuant on top of pruned caches for even greater savings.&lt;/p&gt;
&lt;p&gt;Nobody had validated whether TurboQuant's vector quantization survives the visual token regime. Now someone has.&lt;/p&gt;
&lt;h2&gt;What shipped&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;turboquant-vllm 1.0.0&lt;/strong&gt; is a vLLM plugin, not a fork. It registers via &lt;code&gt;vllm.general_plugins&lt;/code&gt; entry points — the same mechanism vLLM uses for official backends. Install it, pass &lt;code&gt;--attention-backend CUSTOM&lt;/code&gt;, and the TQ4 backend handles everything:&lt;/p&gt;
&lt;p&gt;&lt;img src="https://alberto.codes/turboquant-vllm-architecture.svg" alt="turboquant-vllm plugin architecture" /&gt;&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;&lt;strong&gt;Compress&lt;/strong&gt; — Each new KV vector is rotated by a fixed orthogonal matrix, quantized to 4-bit Lloyd-Max centroids, and nibble-packed (two indices per byte)&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Store&lt;/strong&gt; — Compressed pages use 68 bytes per token per head, vs 256 for FP16&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Decompress&lt;/strong&gt; — Only new tokens are decompressed per decode step (incremental dequantization)&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Attend&lt;/strong&gt; — Standard Flash Attention runs on the decompressed cache&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;For HuggingFace users, &lt;code&gt;CompressedDynamicCache&lt;/code&gt; wraps &lt;code&gt;DynamicCache&lt;/code&gt; and compresses transparently on every &lt;code&gt;cache.update()&lt;/code&gt;.&lt;/p&gt;
&lt;h3&gt;The numbers&lt;/h3&gt;
&lt;p&gt;Molmo2-4B on RTX 4090, 11K visual tokens from a Seinfeld video clip:&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Metric&lt;/th&gt;
&lt;th&gt;Baseline&lt;/th&gt;
&lt;th&gt;TQ4 Compressed&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;KV cache&lt;/td&gt;
&lt;td&gt;1,639 MiB&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;435 MiB (3.76x)&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Output quality&lt;/td&gt;
&lt;td&gt;Detailed scene description&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;Near-identical&lt;/strong&gt; (100+ tokens match word-for-word)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Decode overhead&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;1.78x&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;Molmo2-8B: same 3.76x compression ratio, correctly identifies all Seinfeld characters. Full 23-minute episode processed across 4 clips at 24 tok/s.&lt;/p&gt;
&lt;h2&gt;Design decisions that mattered&lt;/h2&gt;
&lt;h3&gt;Plugin, not fork&lt;/h3&gt;
&lt;p&gt;Other vLLM TurboQuant efforts are forks (brittle, hard to update) or monkey-patches (fragile, version-dependent). &lt;code&gt;turboquant-vllm&lt;/code&gt; uses vLLM's official plugin entry point:&lt;/p&gt;
&lt;pre&gt;&lt;code class="language-toml"&gt;[project.entry-points.&amp;quot;vllm.general_plugins&amp;quot;]
tq4_backend = &amp;quot;turboquant_vllm.vllm:register_tq4_backend&amp;quot;
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;&lt;code&gt;pip install&lt;/code&gt; registers the backend. &lt;code&gt;--attention-backend CUSTOM&lt;/code&gt; activates it. No patching, no forking, no maintenance burden when vLLM updates.&lt;/p&gt;
&lt;h3&gt;Incremental dequantization&lt;/h3&gt;
&lt;p&gt;The naive approach decompresses the entire KV cache at every layer at every decode step. For 11K tokens across 36 layers, that's 3.36x overhead.&lt;/p&gt;
&lt;p&gt;The fix: decompress only the 1 new token per step, append it to a running buffer, let standard Flash Attention handle the rest. Overhead drops to 1.78x. This optimization isn't in the Google paper — it's what makes TQ4 practical for production serving.&lt;/p&gt;
&lt;h3&gt;Cross-platform Triton&lt;/h3&gt;
&lt;p&gt;The fused kernels (compress, decompress, Q@K^T, Flash Attention + TQ4) run on both NVIDIA CUDA and AMD ROCm without code changes. I validated on a Radeon 890M iGPU — 84 of 84 GPU-parametrized tests pass with bit-identical math.&lt;/p&gt;
&lt;p&gt;KV cache compression is most useful on memory-constrained hardware — exactly where AMD's consumer GPUs sit.&lt;/p&gt;
&lt;h2&gt;Validation depth&lt;/h2&gt;
&lt;p&gt;The v1.0.0 release includes:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;180+ tests&lt;/strong&gt; across 9 test files, 95%+ coverage&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;16 GPU experiments&lt;/strong&gt; — each building on the last, documenting failures alongside successes&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Cross-platform validation&lt;/strong&gt; — NVIDIA RTX 4090 + AMD Radeon 890M (ROCm)&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Production container test&lt;/strong&gt; — installed from PyPI into stock &lt;code&gt;vllm/vllm-openai:latest&lt;/code&gt;, served Molmo2-8B video inference with zero errors&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;100% docstring coverage&lt;/strong&gt; enforced by &lt;a href="https://github.com/Alberto-Codes/docvet"&gt;docvet&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;The experiment logs document failure modes nobody else has published: the &lt;a href="https://alberto.codes/blog/2026-03-26-i-ran-turboquant-on-a-vision-model-the-first-output-was-garbage#four-things-i-learned-the-hard-way"&gt;fp16 norms trap&lt;/a&gt; at 10K+ tokens, QJL correction being invisible in standard attention, and multi-layer precision drift in fused kernels. These are landmines in every other implementation that hasn't hit 10K+ visual tokens.&lt;/p&gt;
&lt;h2&gt;What's next&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Upstream vLLM contribution&lt;/strong&gt; — there's an open feature request for TurboQuant support. The plugin is a staging ground.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Flash Attention fusion&lt;/strong&gt; — the fused Triton kernel achieves 17.8x on the Q@K^T micro-benchmark but needs full softmax+V fusion for multi-layer correctness&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Stacking with token pruning&lt;/strong&gt; — combining TurboQuant compression with VL-Cache-style sparsification for multiplicative savings on VLMs&lt;/li&gt;
&lt;/ul&gt;
&lt;blockquote&gt;
&lt;p&gt;Note 2026-09-06: upstream has since shipped TurboQuant in &lt;a href="https://github.com/vllm-project/vllm/pull/38479"&gt;vllm-project/vllm#38479&lt;/a&gt;, and the feature request, &lt;a href="https://github.com/vllm-project/vllm/issues/38201"&gt;vllm-project/vllm#38201&lt;/a&gt;, is closed.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;The full implementation, 16 experiment logs, and architecture docs are at &lt;a href="https://github.com/Alberto-Codes/turboquant-vllm"&gt;github.com/Alberto-Codes/turboquant-vllm&lt;/a&gt;.&lt;/p&gt;
&lt;pre&gt;&lt;code class="language-bash"&gt;pip install turboquant-vllm[vllm]
vllm serve your-model --attention-backend CUSTOM
&lt;/code&gt;&lt;/pre&gt;
</content:encoded>
    </item>
    <item>
      <title>I ran TurboQuant on a vision model. The first output was garbage.</title>
      <link>https://alberto.codes/blog/2026-03-26-i-ran-turboquant-on-a-vision-model-the-first-output-was-garbage</link>
      <guid isPermaLink="true">https://alberto.codes/blog/2026-03-26-i-ran-turboquant-on-a-vision-model-the-first-output-was-garbage</guid>
      <pubDate>Thu, 26 Mar 2026 00:00:00 GMT</pubDate>
      <description>I implemented Google's TurboQuant algorithm for KV cache compression and validated it on Molmo2 video inference on an RTX 4090 — 3.76x compression with near-identical output at 1.78x overhead.</description>
      <content:encoded>&lt;h2&gt;The paper&lt;/h2&gt;
&lt;p&gt;Google published TurboQuant at ICLR 2026 — a method for compressing transformer KV caches from FP16 down to 3-4 bits per coordinate with near-optimal distortion. The theory is elegant: use Lloyd-Max optimal codebooks to quantize each coordinate independently after a random orthogonal rotation that spreads information across dimensions.&lt;/p&gt;
&lt;p&gt;The paper validates on text-only LLMs — Gemma and Mistral. Other implementations have since appeared targeting LLMs on consumer hardware: &lt;a href="https://github.com/tonbistudio/turboquant-pytorch"&gt;turboquant-pytorch&lt;/a&gt; on an RTX 3060, &lt;a href="https://github.com/TheTom/turboquant_plus"&gt;turboquant_plus&lt;/a&gt; on Apple Silicon, and a native &lt;a href="https://github.com/ggml-org/llama.cpp/discussions/20969"&gt;llama.cpp integration&lt;/a&gt;. But as of March 2026, none of them target vision-language models or video input. I wanted to know if TurboQuant works on Molmo2 — Allen AI's open-source VLM — analyzing Seinfeld clips on a single RTX 4090. The TechCrunch coverage hit on March 25. My first commit was 11 PM that night. By 1:33 AM I had 3.76x compression validated on video.&lt;/p&gt;
&lt;h2&gt;Why KV cache matters for video&lt;/h2&gt;
&lt;p&gt;When a vision model processes video, the KV cache grows fast. Molmo2 tokenizes each frame into ~81 visual tokens. A 30-second clip at 2fps produces ~11,000 tokens before the model generates a single word. At FP16, that's 1.6 GB of KV cache across 36 layers.&lt;/p&gt;
&lt;p&gt;On a 24 GB RTX 4090, that 1.6 GB is budget you can't spend on longer clips, larger models, or higher frame rates. Compression directly translates to capability.&lt;/p&gt;
&lt;h2&gt;Experiment 001: garbage&lt;/h2&gt;
&lt;p&gt;I built the full TurboQuant pipeline — Lloyd-Max codebook solver, two-stage quantizer, and a drop-in &lt;code&gt;DynamicCache&lt;/code&gt; wrapper that compresses on write and decompresses on read. The paper recommends TurboQuantProd for keys: 2 bits for MSE quantization, 1 bit for QJL sign correction.&lt;/p&gt;
&lt;p&gt;First test: ask Molmo2-4B &amp;quot;What is 2+2?&amp;quot;&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;Output&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Baseline&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&amp;quot;Four.&amp;quot;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;TurboQuant&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&amp;quot;The number 222 is is a number...&amp;quot; (garbled repetition until max tokens)&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;h2&gt;What the paper assumes you have&lt;/h2&gt;
&lt;p&gt;You have a 3-bit budget per coordinate. The paper's recommended approach, TurboQuantProd, splits it: 2 bits for MSE quantization, 1 bit for QJL sign correction. That split makes sense when you have a custom attention kernel that uses &lt;code&gt;estimate_inner_product()&lt;/code&gt; — QJL correction enables unbiased dot product estimation directly on compressed data.&lt;/p&gt;
&lt;p&gt;But in drop-in mode, standard attention decompresses the keys first and computes &lt;code&gt;Q @ K.T&lt;/code&gt; on the full vectors. QJL is invisible to standard attention. It's like spending a third of your ingredient budget on a garnish that never leaves the kitchen — the diner never sees it, and the dish suffers because you skimped on the protein.&lt;/p&gt;
&lt;p&gt;&lt;img src="https://alberto.codes/turboquant-bit-budget.svg" alt="TurboQuant bit budget: splitting 3 bits into 2-bit MSE plus 1-bit QJL produces 87% cosine similarity and garbled output, while giving all 3 bits to MSE produces 95% cosine similarity and output identical to baseline" /&gt;&lt;/p&gt;
&lt;p&gt;That means TurboQuantProd at 3 bits actually gives you 2-bit MSE reconstruction — roughly 87% cosine similarity. Compounded across 36 layers of autoregressive generation, 87% per layer cascades into noise.&lt;/p&gt;
&lt;p&gt;The fix was one line: switch key compression from TurboQuantProd to TurboQuantMSE. Full 3-bit MSE gives ~95% cosine similarity. The same prompt now returned &amp;quot;Four.&amp;quot; — identical to baseline. On an 11,000-token Seinfeld clip, both baseline and compressed produced coherent scene descriptions with only 1.3x overhead.&lt;/p&gt;
&lt;h2&gt;The compression pipeline&lt;/h2&gt;
&lt;p&gt;&lt;img src="https://alberto.codes/turboquant-kv-compression.svg" alt="TurboQuant compression pipeline: KV tensors at FP16 (256 bytes per token) flow through norm extraction, Haar-random rotation, Lloyd-Max 4-bit scalar quantization, and nibble packing to produce uint8 indices plus fp32 norms at 68 bytes per token — 3.76x compression" /&gt;&lt;/p&gt;
&lt;p&gt;Each KV vector gets its norm extracted (stored as fp32), gets rotated by a shared random orthogonal matrix to spread information across dimensions, then each coordinate is independently quantized using a Lloyd-Max optimal codebook. At 4 bits, two indices pack into one byte — standard bit-shift operations, no custom kernels.&lt;/p&gt;
&lt;h2&gt;Nibble packing: the practical sweet spot&lt;/h2&gt;
&lt;p&gt;3-bit compression stored as uint8 gives only 1.94x — one 3-bit index per byte. Packing across byte boundaries requires custom kernels that don't exist in PyTorch or Triton. 4-bit is trivial: two indices per byte via bit shift. The extra bit also improves reconstruction quality.&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Mode&lt;/th&gt;
&lt;th&gt;Bytes/token&lt;/th&gt;
&lt;th&gt;Compression&lt;/th&gt;
&lt;th&gt;Cosine similarity&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;FP16 baseline&lt;/td&gt;
&lt;td&gt;256&lt;/td&gt;
&lt;td&gt;1.0x&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;TQ3 (unpacked)&lt;/td&gt;
&lt;td&gt;132&lt;/td&gt;
&lt;td&gt;1.94x&lt;/td&gt;
&lt;td&gt;~95%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;TQ4 (nibble-packed)&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;68&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;3.76x&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;~97%&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;TQ3 (bit-packed)&lt;/td&gt;
&lt;td&gt;52&lt;/td&gt;
&lt;td&gt;4.92x&lt;/td&gt;
&lt;td&gt;~95%&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;TQ4 is the sweet spot: 3.76x compression with better quality than TQ3, using only standard PyTorch operations.&lt;/p&gt;
&lt;h2&gt;Cutting the overhead in half&lt;/h2&gt;
&lt;p&gt;My first TQ4 implementation hit 3.76x compression but at 3.36x speed overhead — the cache wrapper re-dequantized all 11,000+ tokens at every layer at every decode step. That's like re-plating the entire service every time a new dish comes off the line. An 88 MB allocation plus a 128x128 rotation matmul, repeated 36 times per generated token.&lt;/p&gt;
&lt;p&gt;The fix: maintain a running decompressed buffer and only dequantize the 1 new token per step — plate the new dish, leave the rest on the pass. Prefill still decompresses everything once, but decode drops from O(seq_len) to O(1) per step.&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Metric&lt;/th&gt;
&lt;th&gt;Full dequant&lt;/th&gt;
&lt;th&gt;Incremental&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Tokens/sec&lt;/td&gt;
&lt;td&gt;8.9&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;16.9&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Overhead&lt;/td&gt;
&lt;td&gt;3.36x&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;1.78x&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Compression&lt;/td&gt;
&lt;td&gt;3.76x&lt;/td&gt;
&lt;td&gt;3.76x&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;The first 100+ tokens match word-for-word between compressed and baseline output.&lt;/p&gt;
&lt;h2&gt;Four things I learned the hard way&lt;/h2&gt;
&lt;ol&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;FP16 norms are a trap.&lt;/strong&gt; At 10K+ tokens across 36 layers, fp16 norm precision loss compounds and flips low-confidence logits. Always use fp32 — it costs 2 extra bytes per vector.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;QJL is invisible in drop-in mode.&lt;/strong&gt; Without a fused attention kernel that operates directly on compressed data, the QJL correction does nothing. Give all bits to MSE.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Peak VRAM is activation-dominated.&lt;/strong&gt; KV cache is ~9% of peak VRAM during prefill. During decode, the ratio shifts as the cache grows — that's where compression pays off. The savings are real in permanent storage but invisible to &lt;code&gt;torch.cuda.max_memory_allocated()&lt;/code&gt; until sequences grow beyond prefill.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Cache your Lloyd-Max codebooks.&lt;/strong&gt; A 36-layer model creates 64 compressors. Without &lt;code&gt;@lru_cache&lt;/code&gt; on the scipy integration, startup takes 2+ minutes.&lt;/p&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;h2&gt;From garbage to 3.76x&lt;/h2&gt;
&lt;p&gt;Two hours earlier, this pipeline produced &amp;quot;The number 222 is is a number...&amp;quot; Now it compresses 1.6 GB of vision-language KV cache into 435 MB with near-identical output quality. The fix that mattered most wasn't an optimization — it was understanding that QJL wastes a bit when you don't have a custom kernel.&lt;/p&gt;
&lt;p&gt;A fused Triton kernel that computes &lt;code&gt;Q @ K.T&lt;/code&gt; directly on compressed data already shows 17.8x speedup on micro-benchmarks. Multi-layer integration is blocked on full Flash Attention-style fusion — computing Q@K^T, softmax, and @V in a single kernel to avoid fp32/bf16 divergence compounding across 36 layers. That's the next milestone.&lt;/p&gt;
&lt;p&gt;The code is MIT-licensed and works with any HuggingFace model:&lt;/p&gt;
&lt;pre&gt;&lt;code class="language-python"&gt;from transformers import DynamicCache
from turboquant_consumer import CompressedDynamicCache

cache = DynamicCache()
compressed = CompressedDynamicCache(cache, head_dim=128, bits=4)  # nibble-packed, 3.76x compression
# Pass cache to model.generate() — compression is transparent
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;The full implementation — 62 tests, five experiment logs, and a benchmark harness — is at &lt;a href="https://github.com/Alberto-Codes/turboquant-consumer"&gt;github.com/Alberto-Codes/turboquant-consumer&lt;/a&gt;.&lt;/p&gt;
</content:encoded>
    </item>
    <item>
      <title>I asked an AI to explain boto3. Then I fixed the docstrings.</title>
      <link>https://alberto.codes/blog/2026-03-23-i-asked-an-ai-to-explain-boto3-then-i-fixed-the-docstrings</link>
      <guid isPermaLink="true">https://alberto.codes/blog/2026-03-23-i-asked-an-ai-to-explain-boto3-then-i-fixed-the-docstrings</guid>
      <pubDate>Mon, 23 Mar 2026 00:00:00 GMT</pubDate>
      <description>I cloned the most downloaded Python package twice, fixed the docstrings with docvet, and asked AI to generate architecture documentation from both. The results weren't even close.</description>
      <content:encoded>&lt;h2&gt;The experiment&lt;/h2&gt;
&lt;p&gt;boto3 is the most downloaded package on PyPI — 43 million installs a day. Every AI coding assistant that helps you write AWS code reads its docstrings. But how good are those docstrings, and does it matter?&lt;/p&gt;
&lt;p&gt;I ran an experiment. Clone boto3 twice at the same commit (&lt;code&gt;04dfc51&lt;/code&gt;, v1.42.73). Leave one copy untouched. Run &lt;a href="https://github.com/Alberto-Codes/docvet"&gt;docvet&lt;/a&gt; on the other — 336 findings across 39 files, 50.3% docstring coverage — and fix every finding. Then ask a fresh AI agent to generate &lt;code&gt;ARCHITECTURE.md&lt;/code&gt; from each copy, with no knowledge of what changed.&lt;/p&gt;
&lt;p&gt;Same codebase. Same model. Same prompt.&lt;/p&gt;
&lt;p&gt;The only difference: docstring quality.&lt;/p&gt;
&lt;h2&gt;What docvet found&lt;/h2&gt;
&lt;p&gt;The audit surfaced gaps across every layer of boto3's documentation:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;147 missing docstrings&lt;/strong&gt; — half of all public symbols had no documentation at all&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;82 missing Returns sections&lt;/strong&gt; — functions that return values but don't say what&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;27 missing Attributes sections&lt;/strong&gt; — classes with undocumented public attributes&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;13 missing Raises sections&lt;/strong&gt; — exceptions thrown but never mentioned in docs&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;45 missing Examples&lt;/strong&gt; — public classes with no usage examples&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;After a second pass with strict configuration (&lt;code&gt;ignore-private = false&lt;/code&gt;, &lt;code&gt;ignore-magic = false&lt;/code&gt;), docvet also surfaced gaps in private methods like &lt;code&gt;_register_default_handlers&lt;/code&gt; — the architectural keystone that wires boto3's entire event-driven customization system.&lt;/p&gt;
&lt;h2&gt;Two architecture documents, one codebase&lt;/h2&gt;
&lt;p&gt;&lt;img src="https://alberto.codes/boto3-architecture-coverage.svg" alt="Architecture coverage comparison: the original agent covered 5 subsystems in 618 lines, while the docvet-fixed agent covered 9 subsystems in 484 lines — including the event system, CRT backend, DynamoDB pipeline, and docs generation" /&gt;&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;Without docvet&lt;/th&gt;
&lt;th&gt;With docvet&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Lines&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;618&lt;/td&gt;
&lt;td&gt;484&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Mermaid diagrams&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;13&lt;/td&gt;
&lt;td&gt;15&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Resource factory&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Covered&lt;/td&gt;
&lt;td&gt;Covered&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Action execution&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Covered&lt;/td&gt;
&lt;td&gt;Covered&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Collection pagination&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Covered&lt;/td&gt;
&lt;td&gt;Covered&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Session lifecycle&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Covered&lt;/td&gt;
&lt;td&gt;Covered&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Event-driven customization&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Not covered&lt;/td&gt;
&lt;td&gt;Full flowchart&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;CRT transfer backend&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Not mentioned&lt;/td&gt;
&lt;td&gt;Dedicated diagram&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;DynamoDB pipeline&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Basic&lt;/td&gt;
&lt;td&gt;3 diagrams deep&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Docs generation system&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Absent&lt;/td&gt;
&lt;td&gt;Flowchart&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Exception hierarchy&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Not covered&lt;/td&gt;
&lt;td&gt;Class diagram&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;The untouched agent went deep on the subsystems it could reverse-engineer from code — resource factory internals, action execution flows, collection pagination. But it missed the subsystems that depend on documentation to discover: the event-driven customization wiring, the CRT backend, the documentation generation pipeline.&lt;/p&gt;
&lt;p&gt;The docvet-fixed agent covered everything the first agent did, plus five additional subsystems — in 134 fewer lines. It didn't need to spend tokens reverse-engineering what the docstrings already explained.&lt;/p&gt;
&lt;p&gt;The difference is clearest in the event system. The agent working with the original docstrings noted: &lt;em&gt;&amp;quot;Session._register_default_handlers() has no docstring at all. This is arguably the most architecturally important method in the codebase. I had to read every register() call and trace the lazy_call targets to understand the customization architecture.&amp;quot;&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;The docvet-fixed agent didn't complain about that method. It diagrammed it — because the docstring told it what was there:&lt;/p&gt;
&lt;p&gt;&lt;img src="https://alberto.codes/boto3-event-system.svg" alt="boto3's event-driven customization system: Session._register_default_handlers() wires S3 transfers, DynamoDB transforms, and EC2 tag injection via botocore events. The original agent couldn't produce this diagram." /&gt;&lt;/p&gt;
&lt;h2&gt;Why docstrings change AI comprehension&lt;/h2&gt;
&lt;p&gt;The agents got the same facts right. Both understood that boto3 wraps botocore, that the resource factory dynamically generates classes from JSON, that collections handle pagination transparently. The difference was in &lt;em&gt;how they got there&lt;/em&gt;.&lt;/p&gt;
&lt;p&gt;Without docstrings, the agent reverse-engineers. It reads function bodies, traces imports, follows call chains. This works — AI models are remarkably good at it — but it's expensive in tokens and narrow in scope. The agent spends its budget understanding &lt;em&gt;how&lt;/em&gt; individual functions work and runs out before mapping &lt;em&gt;how modules connect&lt;/em&gt;.&lt;/p&gt;
&lt;p&gt;With docstrings, the agent comprehends. It reads the module docstring, follows the See Also cross-references, checks the Returns and Raises sections, and moves on. It spends less time on each function and more time on the architecture. The result is broader coverage in fewer lines.&lt;/p&gt;
&lt;p&gt;This aligns with what the research shows. &lt;a href="https://arxiv.org/abs/2404.03114"&gt;Macke &amp;amp; Doyle (NAACL 2024)&lt;/a&gt; found that incorrect documentation degrades LLM task success by 22.6 percentage points — while missing documentation has no statistically significant effect on accuracy. The AI gets the &lt;em&gt;answers&lt;/em&gt; right either way. But the path matters: reverse-engineering is slower, narrower, and misses the connections between modules that docstrings make explicit.&lt;/p&gt;
&lt;h2&gt;The pop quiz&lt;/h2&gt;
&lt;p&gt;I tested this with targeted questions across multiple models. The cleanest example: I asked Sonnet what exceptions &lt;code&gt;S3Transfer.upload_file()&lt;/code&gt; raises and how to handle them.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Without docvet&lt;/strong&gt; — the agent reported: &lt;em&gt;&amp;quot;The docstring says only 'Upload a file to an S3 object' — zero mention of type validation or failure behavior.&amp;quot;&lt;/em&gt; It found the right answer (&lt;code&gt;ValueError&lt;/code&gt; and &lt;code&gt;S3UploadFailedError&lt;/code&gt;) by reading the method body.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;With docvet&lt;/strong&gt; — the agent reported: &lt;em&gt;&amp;quot;The Raises: section names both S3UploadFailedError and ValueError, which is the right starting point.&amp;quot;&lt;/em&gt; Same correct answer, found in the documentation instead of the code.&lt;/p&gt;
&lt;p&gt;I saw this pattern across every model I tested — Opus, Sonnet, Haiku, GPT-4o, GPT-4.1. The answers converged. The sources diverged. In my testing, the docvet-fixed sessions consistently finished faster — the agents spent less time searching code for answers the docstrings already provided.&lt;/p&gt;
&lt;h2&gt;What about wrong docstrings?&lt;/h2&gt;
&lt;p&gt;This experiment only tested missing documentation — I added docstrings where none existed. But the more dangerous case is stale documentation: a docstring that &lt;em&gt;used to be&lt;/em&gt; correct but drifted from the code.&lt;/p&gt;
&lt;p&gt;This is where docvet's freshness checks matter. The &lt;code&gt;stale-signature&lt;/code&gt; rule detects functions whose signatures changed but whose docstrings weren't updated. &lt;code&gt;stale-body&lt;/code&gt; catches implementation changes without corresponding doc updates. The new &lt;code&gt;extra-param-in-docstring&lt;/code&gt; and &lt;code&gt;extra-raises-in-docstring&lt;/code&gt; rules catch docstrings that claim behavior the code no longer exhibits.&lt;/p&gt;
&lt;p&gt;boto3 ships new versions almost daily to track AWS API changes. In a codebase that moves that fast, freshness isn't a nice-to-have — it's the difference between documentation that helps your AI tools and documentation that actively misleads them.&lt;/p&gt;
&lt;h2&gt;What this means for your codebase&lt;/h2&gt;
&lt;p&gt;boto3 is maintained by Amazon. It has 50.3% docstring coverage. If the most downloaded Python package has 336 documentation gaps that affect how AI understands its architecture, your codebase almost certainly has more.&lt;/p&gt;
&lt;p&gt;The fix isn't writing docstrings for the sake of coverage metrics. It's writing docstrings that tell AI agents what they need to know: what a function returns, what it raises, how modules connect, and where to look next. docvet identifies exactly where those gaps are — and with the right configuration, it catches them in private methods and magic methods too.&lt;/p&gt;
&lt;pre&gt;&lt;code class="language-bash"&gt;pip install docvet
# or
uv add docvet --dev

docvet check --all --verbose
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Full documentation: &lt;a href="https://alberto-codes.github.io/docvet/"&gt;alberto-codes.github.io/docvet&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;The AI reading your code will find the answers either way. The question is whether it finds them in your documentation or reverse-engineers them from your implementation. One path is faster, broader, and produces better results. docvet makes sure the documentation is there when the AI comes looking.&lt;/p&gt;
&lt;p&gt;Both architecture documents — &lt;a href="https://gist.github.com/Alberto-Codes/86bef0f945d6d8e30c31dbb7478c2d6e"&gt;original and docvet-fixed&lt;/a&gt; — are available for anyone who wants to compare them side by side.&lt;/p&gt;
</content:encoded>
    </item>
    <item>
      <title>When docstrings lie, your AI tools pay the price</title>
      <link>https://alberto.codes/blog/2026-03-22-when-docstrings-lie-your-ai-tools-pay-the-price</link>
      <guid isPermaLink="true">https://alberto.codes/blog/2026-03-22-when-docstrings-lie-your-ai-tools-pay-the-price</guid>
      <pubDate>Sun, 22 Mar 2026 00:00:00 GMT</pubDate>
      <description>Wrong documentation hurts AI tools more than missing documentation. docvet 1.14 introduces bidirectional verification — checking both what your docstrings fail to mention and what they wrongly claim.</description>
      <content:encoded>&lt;h2&gt;The problem isn't missing docstrings&lt;/h2&gt;
&lt;p&gt;Most docstring linters answer one question: does this function have a docstring? That's useful, but it's the wrong question for 2026.&lt;/p&gt;
&lt;p&gt;A &lt;a href="https://arxiv.org/abs/2404.03114"&gt;2024 study by Macke &amp;amp; Doyle&lt;/a&gt; found that &lt;strong&gt;incorrect documentation degrades LLM task success by 22.6 percentage points&lt;/strong&gt; — while missing documentation has no statistically significant effect. Read that again. Your AI coding assistant performs &lt;em&gt;worse&lt;/em&gt; when your docstrings are wrong than when they don't exist at all.&lt;/p&gt;
&lt;p&gt;The reason is straightforward. When a docstring says a function accepts &lt;code&gt;timeout&lt;/code&gt; but the parameter was renamed to &lt;code&gt;max_wait&lt;/code&gt; three commits ago, every tool that reads that docstring — Copilot, Claude, your IDE's autocomplete — generates code that passes an argument the function doesn't accept. Missing docs let the model fall back on reading the code. Wrong docs actively mislead it.&lt;/p&gt;
&lt;p&gt;docvet 1.14 is built around this insight. Instead of just checking whether docstrings &lt;em&gt;exist&lt;/em&gt;, it checks whether they &lt;em&gt;match the code&lt;/em&gt;.&lt;/p&gt;
&lt;h2&gt;What bidirectional verification means&lt;/h2&gt;
&lt;p&gt;&lt;img src="https://alberto.codes/bidirectional-verification.svg" alt="Bidirectional verification: forward checks catch missing documentation, reverse checks catch documented behavior that doesn't exist in code" /&gt;&lt;/p&gt;
&lt;p&gt;Before 1.14, docvet's enrichment checks worked in one direction: they looked at the code and asked &amp;quot;did the docstring mention this?&amp;quot; If your function raised &lt;code&gt;ValueError&lt;/code&gt; but the docstring had no &lt;code&gt;Raises:&lt;/code&gt; section, docvet flagged it. Useful — but only half the picture.&lt;/p&gt;
&lt;p&gt;The other half is the reverse question: does the docstring &lt;em&gt;claim&lt;/em&gt; something the code doesn't actually do?&lt;/p&gt;
&lt;p&gt;A docstring that says a function raises &lt;code&gt;FileNotFoundError&lt;/code&gt; when the function never raises anything isn't just stale — it's a trap. Callers write &lt;code&gt;try/except FileNotFoundError&lt;/code&gt; blocks that never trigger, or worse, they restructure their error handling around phantom exceptions. AI tools generate defensive code for errors that can't happen.&lt;/p&gt;
&lt;p&gt;v1.14 adds three reverse enrichment rules:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;extra-raises-in-docstring&lt;/strong&gt; — documents exceptions the function doesn't raise&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;extra-yields-in-docstring&lt;/strong&gt; — documents yields in a non-generator function&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;extra-returns-in-docstring&lt;/strong&gt; — documents return values the function never returns&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Combined with the existing forward checks, docvet now performs full bidirectional verification: every claim in the docstring has a corresponding behavior in the code, and every behavior in the code has a corresponding claim in the docstring. With this release, docvet reaches 31 rules across six quality layers — and truthfulness is now the central theme.&lt;/p&gt;
&lt;h2&gt;Parameter agreement: where drift actually lives&lt;/h2&gt;
&lt;p&gt;&lt;img src="https://alberto.codes/param-drift-lifecycle.svg" alt="Parameter drift lifecycle: a refactor renames a parameter, call sites are updated, but the docstring is forgotten — AI reads the stale docstring and generates wrong code. docvet intercepts at the forgotten step." /&gt;&lt;/p&gt;
&lt;p&gt;The most impactful new feature in 1.14 is parameter agreement checking. Two rules — &lt;code&gt;missing-param-in-docstring&lt;/code&gt; and &lt;code&gt;extra-param-in-docstring&lt;/code&gt; — compare the &lt;code&gt;Args:&lt;/code&gt; section against the function signature, parameter by parameter.&lt;/p&gt;
&lt;p&gt;This catches the single most common form of documentation drift: renamed, added, or removed parameters where the docstring wasn't updated. It happens constantly in active codebases. You rename &lt;code&gt;retries&lt;/code&gt; to &lt;code&gt;max_retries&lt;/code&gt; across a refactor, update every call site, and forget the one place that still says &lt;code&gt;retries&lt;/code&gt; — the docstring.&lt;/p&gt;
&lt;p&gt;The checks handle real-world Python signatures: positional-only parameters (PEP 570), keyword-only parameters, &lt;code&gt;self&lt;/code&gt;/&lt;code&gt;cls&lt;/code&gt; exclusion, and optional &lt;code&gt;*args&lt;/code&gt;/&lt;code&gt;**kwargs&lt;/code&gt; filtering. Both Google and Sphinx docstring styles are supported.&lt;/p&gt;
&lt;p&gt;Here's what a finding looks like:&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;src/client.py:47: missing-param-in-docstring Function 'connect' has parameters not documented in Args: max_retries [required]
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;That's a one-line signal that would otherwise become a user-filed bug report, or worse, silently wrong AI-generated code.&lt;/p&gt;
&lt;h2&gt;Catching docstrings that say nothing&lt;/h2&gt;
&lt;p&gt;Not all low-quality docstrings are wrong — some are just empty calories. The new &lt;code&gt;trivial-docstring&lt;/code&gt; rule detects summary lines that restate the symbol name without adding information:&lt;/p&gt;
&lt;pre&gt;&lt;code class="language-python"&gt;def get_user():
    &amp;quot;&amp;quot;&amp;quot;Get user.&amp;quot;&amp;quot;&amp;quot;
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;This docstring technically exists. It passes every presence check. But it tells you nothing you couldn't read from the function name. docvet decomposes both the symbol name and the summary into word sets (handling CamelCase and snake_case), filters stop words, and flags cases where the summary is a subset of the name.&lt;/p&gt;
&lt;p&gt;The rule skips &lt;code&gt;@property&lt;/code&gt; and &lt;code&gt;@cached_property&lt;/code&gt; — those frequently have legitimate one-line summaries that mirror the attribute name.&lt;/p&gt;
&lt;h2&gt;Deprecation and constructor documentation&lt;/h2&gt;
&lt;p&gt;Two more rules round out the release:&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;missing-deprecation&lt;/strong&gt; detects functions that use &lt;code&gt;warnings.warn(DeprecationWarning)&lt;/code&gt; or the &lt;code&gt;@deprecated&lt;/code&gt; decorator (PEP 702) without mentioning deprecation anywhere in their docstring. The match is intentionally loose — the word &amp;quot;deprecated&amp;quot; anywhere in the docstring satisfies it, because there's no single standard format for deprecation notices.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;undocumented-init-params&lt;/strong&gt; catches classes whose &lt;code&gt;__init__&lt;/code&gt; accepts parameters but neither the class docstring nor the &lt;code&gt;__init__&lt;/code&gt; docstring has an &lt;code&gt;Args:&lt;/code&gt; section. This one defaults to off — it's opt-in for teams that want it.&lt;/p&gt;
&lt;h2&gt;Design tradeoffs&lt;/h2&gt;
&lt;p&gt;A few deliberate choices worth calling out:&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Reverse checks use &lt;code&gt;recommended&lt;/code&gt; severity, not &lt;code&gt;required&lt;/code&gt;.&lt;/strong&gt; Forward checks (&amp;quot;you raise but don't document it&amp;quot;) have low false-positive rates because the code is the source of truth. Reverse checks (&amp;quot;you document but don't raise it&amp;quot;) have higher false-positive risk — a function might delegate to a helper that raises, or a parent class might document exceptions from subclass overrides. We chose &lt;code&gt;recommended&lt;/code&gt; to surface these without blocking CI pipelines.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Two rules default to off.&lt;/strong&gt; &lt;code&gt;missing-return-type&lt;/code&gt; and &lt;code&gt;undocumented-init-params&lt;/code&gt; are opt-in. Not every team wants to enforce return types in docstrings when they already have type annotations, and not every class needs &lt;code&gt;__init__&lt;/code&gt; parameter docs in the class body. Progressive adoption matters.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Parameter agreement checks support both directions simultaneously.&lt;/strong&gt; Rather than shipping &amp;quot;check if params are documented&amp;quot; and &amp;quot;check if documented params exist&amp;quot; as a single rule, they're separate rules gated by the same &lt;code&gt;require-param-agreement&lt;/code&gt; config key. This means you can see exactly which direction the drift went — missing documentation vs. stale documentation — in your CI output.&lt;/p&gt;
&lt;h2&gt;Who this is for&lt;/h2&gt;
&lt;p&gt;If you maintain a Python library that other developers (or AI tools) consume, v1.14 catches the class of documentation bugs that cause the most downstream damage. Parameter mismatches, phantom exceptions, and trivial summaries are the documentation equivalent of type errors — they compile fine but break at runtime.&lt;/p&gt;
&lt;p&gt;If you're using docvet in CI already, upgrade and enable the new checks:&lt;/p&gt;
&lt;pre&gt;&lt;code class="language-bash"&gt;pip install docvet
# or
uv add docvet --dev
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;The param agreement checks are on by default. Reverse enrichment checks are on by default. &lt;code&gt;missing-return-type&lt;/code&gt; and &lt;code&gt;undocumented-init-params&lt;/code&gt; are opt-in — enable them in &lt;code&gt;pyproject.toml&lt;/code&gt; when you're ready:&lt;/p&gt;
&lt;pre&gt;&lt;code class="language-toml"&gt;[tool.docvet.enrichment]
require-return-type = true
require-init-params = true
&lt;/code&gt;&lt;/pre&gt;
&lt;h2&gt;What's next&lt;/h2&gt;
&lt;p&gt;v1.14 brings docvet to 31 rules across six quality layers. The next focus area is expanding what &amp;quot;truthfulness&amp;quot; means — moving beyond structural checks into semantic verification. The question isn't just &amp;quot;did you document the parameters?&amp;quot; but &amp;quot;is what you said about them accurate?&amp;quot;&lt;/p&gt;
&lt;p&gt;Every function with a stale &lt;code&gt;Args:&lt;/code&gt; section is a wrong answer waiting to happen — in your IDE, in your code review, in the pull request your AI assistant is about to generate. docvet 1.14 catches them before they ship.&lt;/p&gt;
</content:encoded>
    </item>
    <item>
      <title>Encrypt ADK Sessions in 5 Minutes</title>
      <link>https://alberto.codes/blog/2026-03-06-encrypt-adk-sessions-in-five-minutes</link>
      <guid isPermaLink="true">https://alberto.codes/blog/2026-03-06-encrypt-adk-sessions-in-five-minutes</guid>
      <pubDate>Fri, 06 Mar 2026 00:00:00 GMT</pubDate>
      <description>Install adk-secure-sessions, swap one import, and verify your agent's session data is encrypted at rest — start to finish in under 5 minutes.</description>
      <content:encoded>&lt;p&gt;&lt;a href="https://alberto.codes/blog/2026-03-01-your-ai-agents-memories-arent-encrypted"&gt;In Part 1&lt;/a&gt;, I showed that Google ADK stores everything your agent knows — tool calls, user messages, conversation context — in plaintext SQLite. If that made you uncomfortable, this post fixes it.&lt;/p&gt;
&lt;p&gt;This is the recipe card. Ingredients, steps, done. The explanation of why the soufflé rises comes later in Part 3.&lt;/p&gt;
&lt;hr /&gt;
&lt;h2&gt;Prerequisites&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Python 3.12+&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;An existing ADK agent&lt;/strong&gt; using &lt;code&gt;DatabaseSessionService&lt;/code&gt; (or a willingness to create a minimal one)&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;No system libraries, no C compilation, no Docker. The library is pure Python with two runtime dependencies: &lt;code&gt;google-adk&lt;/code&gt; and &lt;code&gt;cryptography&lt;/code&gt;. A short ingredient list.&lt;/p&gt;
&lt;hr /&gt;
&lt;h2&gt;Step 1: Install&lt;/h2&gt;
&lt;pre&gt;&lt;code class="language-bash"&gt;pip install adk-secure-sessions
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Or with &lt;a href="https://docs.astral.sh/uv/"&gt;uv&lt;/a&gt;:&lt;/p&gt;
&lt;pre&gt;&lt;code class="language-bash"&gt;uv add adk-secure-sessions
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Verify the install:&lt;/p&gt;
&lt;pre&gt;&lt;code class="language-bash"&gt;python -c &amp;quot;import adk_secure_sessions; print('OK')&amp;quot;
&lt;/code&gt;&lt;/pre&gt;
&lt;hr /&gt;
&lt;h2&gt;Step 2: Swap the Import&lt;/h2&gt;
&lt;p&gt;Your agent code probably has something like this:&lt;/p&gt;
&lt;pre&gt;&lt;code class="language-python"&gt;# Before — ADK default (unencrypted):
from google.adk.sessions import DatabaseSessionService

session_service = DatabaseSessionService(
    db_url=&amp;quot;sqlite+aiosqlite:///sessions.db&amp;quot;
)
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Replace it with:&lt;/p&gt;
&lt;pre&gt;&lt;code class="language-python"&gt;# After — encrypted:
from adk_secure_sessions import EncryptedSessionService, FernetBackend

session_service = EncryptedSessionService(
    db_url=&amp;quot;sqlite+aiosqlite:///sessions.db&amp;quot;,
    backend=FernetBackend(&amp;quot;your-secret-passphrase&amp;quot;),
)
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Two changes: the import line and the constructor. Everything else in your agent stays the same — &lt;code&gt;create_session&lt;/code&gt;, &lt;code&gt;get_session&lt;/code&gt;, &lt;code&gt;list_sessions&lt;/code&gt;, &lt;code&gt;delete_session&lt;/code&gt;, &lt;code&gt;append_event&lt;/code&gt; — the full ADK session lifecycle, identical behavior. The difference is what hits the disk.&lt;/p&gt;
&lt;hr /&gt;
&lt;h2&gt;Step 3: Use the Async Context Manager&lt;/h2&gt;
&lt;p&gt;For proper connection cleanup, wrap the service in &lt;code&gt;async with&lt;/code&gt;. Here's a complete, runnable script:&lt;/p&gt;
&lt;pre&gt;&lt;code class="language-python"&gt;import asyncio
from adk_secure_sessions import EncryptedSessionService, FernetBackend


async def main():
    backend = FernetBackend(&amp;quot;my-secret-passphrase&amp;quot;)

    async with EncryptedSessionService(
        db_url=&amp;quot;sqlite+aiosqlite:///sessions.db&amp;quot;,
        backend=backend,
    ) as service:
        # Create a session with sensitive state
        session = await service.create_session(
            app_name=&amp;quot;my-agent&amp;quot;,
            user_id=&amp;quot;user-123&amp;quot;,
            state={
                &amp;quot;patient_name&amp;quot;: &amp;quot;Jane Doe&amp;quot;,
                &amp;quot;diagnosis_code&amp;quot;: &amp;quot;J06.9&amp;quot;,
                &amp;quot;api_key&amp;quot;: &amp;quot;sk-secret-key-12345&amp;quot;,
            },
        )
        print(f&amp;quot;Created session: {session.id}&amp;quot;)

        # Retrieve — state is automatically decrypted
        session = await service.get_session(
            app_name=&amp;quot;my-agent&amp;quot;,
            user_id=&amp;quot;user-123&amp;quot;,
            session_id=session.id,
        )
        print(f&amp;quot;Decrypted state: {session.state}&amp;quot;)

        # List sessions for this app/user
        response = await service.list_sessions(
            app_name=&amp;quot;my-agent&amp;quot;,
            user_id=&amp;quot;user-123&amp;quot;,
        )
        print(f&amp;quot;Sessions found: {len(response.sessions)}&amp;quot;)

        # Clean up when you're done
        await service.delete_session(
            app_name=&amp;quot;my-agent&amp;quot;,
            user_id=&amp;quot;user-123&amp;quot;,
            session_id=session.id,
        )
        print(&amp;quot;Session deleted&amp;quot;)


asyncio.run(main())
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Copy this into a file and run it. The API behaves identically to ADK's &lt;code&gt;DatabaseSessionService&lt;/code&gt; — same methods, same signatures, same return types. The only difference is what's stored on disk: switching from a glass jar to a lockbox. Same ingredients go in, same ingredients come out, but nobody can peek inside without the key.&lt;/p&gt;
&lt;hr /&gt;
&lt;h2&gt;Step 4: Verify the Encryption&lt;/h2&gt;
&lt;p&gt;Trust but verify. Open the SQLite database directly and confirm the data is actually encrypted.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Using the &lt;code&gt;sqlite3&lt;/code&gt; CLI:&lt;/strong&gt;&lt;/p&gt;
&lt;pre&gt;&lt;code class="language-bash"&gt;sqlite3 sessions.db &amp;quot;SELECT state FROM sessions LIMIT 1;&amp;quot;
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;You'll see a base64-encoded string — the encrypted envelope — not readable JSON:&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;AQFnQUFBQUJuVm1Gc2RX...
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;&lt;strong&gt;Using Python:&lt;/strong&gt;&lt;/p&gt;
&lt;pre&gt;&lt;code class="language-python"&gt;import sqlite3

conn = sqlite3.connect(&amp;quot;sessions.db&amp;quot;)
row = conn.execute(&amp;quot;SELECT state FROM sessions LIMIT 1&amp;quot;).fetchone()
print(row[0][:60])  # First 60 chars of the encrypted envelope
conn.close()
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;What you won't see: &lt;code&gt;{&amp;quot;patient_name&amp;quot;: &amp;quot;Jane Doe&amp;quot;, &amp;quot;diagnosis_code&amp;quot;: &amp;quot;J06.9&amp;quot;}&lt;/code&gt;. That's the point. With &lt;code&gt;DatabaseSessionService&lt;/code&gt;, anyone with file access reads your mise en place. With &lt;code&gt;EncryptedSessionService&lt;/code&gt;, they see noise.&lt;/p&gt;
&lt;p&gt;For a more convincing demo, run the &lt;a href="https://github.com/Alberto-Codes/adk-secure-sessions/blob/main/examples/basic_usage.py"&gt;basic usage example&lt;/a&gt; from the repo — it runs a real multi-turn ADK agent with Ollama and then inspects the raw database to prove no plaintext leaks. After a three-turn conversation about patient intake, the database contains zero occurrences of &amp;quot;Jane Doe&amp;quot; or &amp;quot;headache.&amp;quot;&lt;/p&gt;
&lt;hr /&gt;
&lt;h2&gt;Step 5: Manage Your Passphrase&lt;/h2&gt;
&lt;p&gt;The passphrase is the only secret. Never hardcode it.&lt;/p&gt;
&lt;pre&gt;&lt;code class="language-python"&gt;import os
from adk_secure_sessions import EncryptedSessionService, FernetBackend

backend = FernetBackend(os.environ[&amp;quot;SESSION_KEY&amp;quot;])
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Set it in your environment, your &lt;code&gt;.env&lt;/code&gt; file, or your secrets manager. The library handles everything else — &lt;code&gt;FernetBackend&lt;/code&gt; derives a cryptographic key using PBKDF2-HMAC-SHA256 with 480,000 iterations. You don't need to generate, store, or rotate raw key material.&lt;/p&gt;
&lt;p&gt;If you use the wrong passphrase to read a session encrypted with a different one, you get a clear &lt;code&gt;DecryptionError&lt;/code&gt; — never garbage data, never silent corruption.&lt;/p&gt;
&lt;hr /&gt;
&lt;h2&gt;What You Just Built&lt;/h2&gt;
&lt;p&gt;Five steps, plaintext to encrypted-at-rest:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;&lt;strong&gt;Installed&lt;/strong&gt; — &lt;code&gt;pip install adk-secure-sessions&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Swapped&lt;/strong&gt; — one import, one constructor change&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Ran&lt;/strong&gt; — same API, encrypted storage&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Verified&lt;/strong&gt; — the database contains ciphertext, not JSON&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Secured&lt;/strong&gt; — passphrase in the environment, not the codebase&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;Your agent still works the same way. Your tests still pass. But the SQLite file is now useless without the key — like a walk-in freezer with a combination lock. Nothing changes about how the food is stored or retrieved, but the back door isn't open anymore.&lt;/p&gt;
&lt;hr /&gt;
&lt;h2&gt;Error Handling&lt;/h2&gt;
&lt;p&gt;When things go wrong, the library tells you what happened:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;&lt;code&gt;ConfigurationError&lt;/code&gt;&lt;/strong&gt; — raised at startup if the backend is misconfigured. You'll catch this before any data is written.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;&lt;code&gt;DecryptionError&lt;/code&gt;&lt;/strong&gt; — raised if you read a session with the wrong key. The library never returns garbage.&lt;/li&gt;
&lt;/ul&gt;
&lt;pre&gt;&lt;code class="language-python"&gt;from adk_secure_sessions import (
    ConfigurationError,
    DecryptionError,
    EncryptedSessionService,
    FernetBackend,
)

try:
    async with EncryptedSessionService(
        db_url=&amp;quot;sqlite+aiosqlite:///sessions.db&amp;quot;,
        backend=FernetBackend(&amp;quot;correct-passphrase&amp;quot;),
    ) as service:
        session = await service.get_session(
            app_name=&amp;quot;my-agent&amp;quot;,
            user_id=&amp;quot;user-123&amp;quot;,
            session_id=&amp;quot;some-session-id&amp;quot;,
        )
        if session is None:
            print(&amp;quot;Session not found&amp;quot;)
except ConfigurationError:
    print(&amp;quot;Backend doesn't conform to EncryptionBackend protocol&amp;quot;)
except DecryptionError:
    print(&amp;quot;Wrong key — cannot decrypt session data&amp;quot;)
&lt;/code&gt;&lt;/pre&gt;
&lt;hr /&gt;
&lt;h2&gt;Key Takeaways&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;One install, one import change&lt;/strong&gt; — &lt;code&gt;pip install adk-secure-sessions&lt;/code&gt;, swap the constructor, done&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Full ADK lifecycle&lt;/strong&gt; — &lt;code&gt;create_session&lt;/code&gt;, &lt;code&gt;get_session&lt;/code&gt;, &lt;code&gt;list_sessions&lt;/code&gt;, &lt;code&gt;delete_session&lt;/code&gt;, and &lt;code&gt;append_event&lt;/code&gt; all work identically&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Verify it yourself&lt;/strong&gt; — inspect the SQLite file to confirm ciphertext, not plaintext&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Passphrase management&lt;/strong&gt; — use environment variables, never hardcode secrets&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Clear errors&lt;/strong&gt; — &lt;code&gt;DecryptionError&lt;/code&gt; for wrong keys, &lt;code&gt;ConfigurationError&lt;/code&gt; for bad setup&lt;/li&gt;
&lt;/ul&gt;
&lt;hr /&gt;
&lt;h2&gt;Further Reading&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href="https://alberto-codes.github.io/adk-secure-sessions/"&gt;adk-secure-sessions documentation&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://alberto-codes.github.io/adk-secure-sessions/getting-started/"&gt;Getting Started guide&lt;/a&gt; — the library's own tutorial with additional details&lt;/li&gt;
&lt;li&gt;&lt;a href="https://pypi.org/project/adk-secure-sessions/"&gt;adk-secure-sessions on PyPI&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://google.github.io/adk-docs/sessions/session/"&gt;Google ADK Session docs&lt;/a&gt; — the underlying session model this library encrypts&lt;/li&gt;
&lt;li&gt;Previous in this series: &lt;a href="https://alberto.codes/blog/2026-03-01-your-ai-agents-memories-arent-encrypted"&gt;Your AI Agent's Memories Aren't Encrypted&lt;/a&gt; (Part 1)&lt;/li&gt;
&lt;li&gt;Next in this series: The Architecture of Encrypted Sessions (Part 3)&lt;/li&gt;
&lt;/ul&gt;
</content:encoded>
    </item>
    <item>
      <title>Stop Writing AI Agent Prompts by Hand</title>
      <link>https://alberto.codes/blog/2026-03-05-stop-writing-ai-agent-prompts-by-hand</link>
      <guid isPermaLink="true">https://alberto.codes/blog/2026-03-05-stop-writing-ai-agent-prompts-by-hand</guid>
      <pubDate>Thu, 05 Mar 2026 00:00:00 GMT</pubDate>
      <description>You write your agent's instructions, test them, tweak a word, test again, and hope the change helped. There's an algorithm that does this better than you do — evolutionary optimization finds prompts you'd never write yourself.</description>
      <content:encoded>&lt;p&gt;You write a prompt. You test it. It's okay but not great — the agent misses edge cases, ignores context, or produces output that's technically correct but tonally wrong. So you tweak a sentence. Test again. Better on one example, worse on another. You rewrite the whole thing. Test again. Repeat until you're tired or the deadline hits, whichever comes first.&lt;/p&gt;
&lt;p&gt;This is how most people build AI agents. It's also how most people season food before they learn to taste as they go — a pinch of this, a dash of that, hope for the best. It works, kind of. But it doesn't scale, it doesn't generalize, and it leaves performance on the table that you'll never find through manual iteration.&lt;/p&gt;
&lt;p&gt;What if an algorithm could do this loop for you — and do it better?&lt;/p&gt;
&lt;hr /&gt;
&lt;h2&gt;The Problem With Human Prompt Engineering&lt;/h2&gt;
&lt;p&gt;Manual prompt engineering has a ceiling. Humans are good at writing instructions that sound clear to other humans, but LLMs don't process text the way we do. The prompt that reads best to you isn't necessarily the prompt that produces the best output from the model.&lt;/p&gt;
&lt;p&gt;There are a few specific failure modes:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Local optima&lt;/strong&gt; — you find a prompt that works well enough and stop iterating, never discovering the dramatically different phrasing that scores higher&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Narrow testing&lt;/strong&gt; — you test against two or three examples, miss the edge cases, and discover them in production&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Instinct over data&lt;/strong&gt; — you change what feels wrong instead of what measurably underperforms&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Single-objective thinking&lt;/strong&gt; — you optimize for one quality (accuracy, tone, format) while degrading others you're not watching&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;These aren't character flaws. They're the natural limitations of a human trying to search a vast space of possible text through trial and error. The space of possible instructions for even a simple agent is effectively infinite, and your ability to explore it is limited by time, patience, and cognitive bias.&lt;/p&gt;
&lt;hr /&gt;
&lt;h2&gt;Evolutionary Optimization: The Core Idea&lt;/h2&gt;
&lt;p&gt;Evolutionary prompt optimization borrows from genetic algorithms. Instead of you guessing at better prompts, an algorithm systematically explores the space:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;&lt;strong&gt;Evaluate&lt;/strong&gt; — run the agent on a batch of examples, score the outputs&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Reflect&lt;/strong&gt; — analyze what worked and what didn't across all examples&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Mutate&lt;/strong&gt; — an LLM proposes improved instructions based on the analysis&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Select&lt;/strong&gt; — if the new instruction scores better, keep it; otherwise, discard it&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;Repeat until the scores plateau or you hit an iteration limit.&lt;/p&gt;
&lt;p&gt;&lt;img src="https://alberto.codes/evolution-loop.svg" alt="The evolution loop — training examples flow through the agent, critic, and reflection agents. The mutated instruction is accepted or rejected based on score improvement, then the loop repeats." /&gt;&lt;/p&gt;
&lt;p&gt;The key insight is step 3. The mutation isn't random — it's informed. A reflection model sees the agent's actual outputs, the scores, and the feedback, then proposes specific text changes to address the weaknesses it observed. It's less &amp;quot;random mutation&amp;quot; and more &amp;quot;systematic recipe refinement&amp;quot; — tasting every dish that comes out of the kitchen, identifying exactly what's off, and adjusting the technique accordingly.&lt;/p&gt;
&lt;p&gt;This is the approach described in the &lt;a href="https://arxiv.org/abs/2507.19457"&gt;GEPA paper&lt;/a&gt; (Genetic-Pareto prompt optimizer). GEPA treats prompt components as evolvable genes and uses LLM-powered reflection instead of random crossover. The result is an optimizer that can improve agent instructions in ways that surprise even the person who wrote the original prompt.&lt;/p&gt;
&lt;hr /&gt;
&lt;h2&gt;What This Looks Like in Practice&lt;/h2&gt;
&lt;p&gt;&lt;a href="https://github.com/Alberto-Codes/gepa-adk"&gt;gepa-adk&lt;/a&gt; is a Python library that brings evolutionary optimization to Google ADK agents. You give it an agent, training examples, and a critic — it gives you back a better prompt.&lt;/p&gt;
&lt;p&gt;Here's the simplest possible example. Start with a greeting agent that has a generic instruction:&lt;/p&gt;
&lt;pre&gt;&lt;code class="language-python"&gt;from google.adk.agents import LlmAgent
from gepa_adk import evolve_sync, EvolutionConfig, SimpleCriticOutput

# The agent to evolve — starts with a vague instruction
agent = LlmAgent(
    name=&amp;quot;greeter&amp;quot;,
    model=&amp;quot;gemini-2.5-flash&amp;quot;,
    instruction=&amp;quot;Greet the user appropriately.&amp;quot;,
)

# A critic that knows what &amp;quot;good&amp;quot; looks like
critic = LlmAgent(
    name=&amp;quot;critic&amp;quot;,
    model=&amp;quot;gemini-2.5-flash&amp;quot;,
    instruction=&amp;quot;Score for formal, Dickens-style greetings. 0.0-1.0.&amp;quot;,
    output_schema=SimpleCriticOutput,
)

# Training examples covering different social contexts
trainset = [
    {&amp;quot;input&amp;quot;: &amp;quot;I am His Majesty, the King.&amp;quot;},
    {&amp;quot;input&amp;quot;: &amp;quot;I am your mother.&amp;quot;},
    {&amp;quot;input&amp;quot;: &amp;quot;I am a close friend.&amp;quot;},
]

result = evolve_sync(agent, trainset, critic=critic)
print(f&amp;quot;Score: {result.original_score:.2f} -&amp;gt; {result.final_score:.2f}&amp;quot;)
print(result.evolved_components[&amp;quot;instruction&amp;quot;])
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;The original instruction — &amp;quot;Greet the user appropriately&amp;quot; — is five words. The evolved instruction might be three paragraphs of specific guidance about formality levels, period-appropriate language, honorific handling, and tonal variation. You would never write that prompt yourself. Not because you couldn't, but because you wouldn't think to be that specific in those particular ways.&lt;/p&gt;
&lt;hr /&gt;
&lt;h2&gt;Why a Critic Changes Everything&lt;/h2&gt;
&lt;p&gt;The critic is what separates evolutionary optimization from &amp;quot;just running the prompt a bunch of times.&amp;quot; Without a critic, you're scoring outputs manually or relying on the agent to self-assess (which is like asking the cook to rate their own soup — helpful but biased).&lt;/p&gt;
&lt;p&gt;A critic agent is a separate LLM that evaluates outputs against explicit criteria. It returns a score and feedback:&lt;/p&gt;
&lt;pre&gt;&lt;code class="language-python"&gt;class SimpleCriticOutput(BaseModel):
    score: float = Field(ge=0.0, le=1.0)
    feedback: str
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;The feedback is the ingredient that makes reflection work. The reflection model doesn't just see &amp;quot;this scored 0.4&amp;quot; — it sees &amp;quot;this scored 0.4 because the greeting was too casual for royalty and used modern slang.&amp;quot; That specific diagnosis drives specific mutations.&lt;/p&gt;
&lt;p&gt;You can also use multi-dimensional scoring — rate outputs on clarity, accuracy, tone, and format independently. GEPA tracks a Pareto frontier across dimensions, so you don't have to collapse everything into a single number and lose information about what's actually improving.&lt;/p&gt;
&lt;hr /&gt;
&lt;h2&gt;The Evolution Loop, Unpacked&lt;/h2&gt;
&lt;p&gt;What actually happens inside &lt;code&gt;evolve()&lt;/code&gt;:&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Iteration 1&lt;/strong&gt;: The agent runs on all training examples with its original instruction. The critic scores each output. Average score: 0.35. The reflection model reads every output, every score, every piece of feedback. It proposes a new instruction that addresses the most common failure patterns.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Iteration 2&lt;/strong&gt;: The agent runs again with the proposed instruction. Scores improve to 0.62. The reflection model sees what got better and what's still weak. It proposes another refinement.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Iteration 3&lt;/strong&gt;: Scores hit 0.78. The reflection model notices diminishing returns and makes a smaller, more targeted adjustment.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Iteration 4&lt;/strong&gt;: Scores reach 0.81. No improvement over the previous best after the patience window. Evolution stops.&lt;/p&gt;
&lt;p&gt;The &lt;code&gt;EvolutionConfig&lt;/code&gt; controls this process:&lt;/p&gt;
&lt;pre&gt;&lt;code class="language-python"&gt;config = EvolutionConfig(
    max_iterations=5,     # Upper bound on iterations
    patience=2,           # Stop if no improvement for N iterations
    reflection_model=&amp;quot;gemini-2.5-flash&amp;quot;,  # Model for generating mutations
)
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Each iteration makes multiple LLM calls — the agent runs on every training example, the critic scores every output, and the reflection model analyzes everything. This is why gepa-adk recommends local models via Ollama for development. Evolution is compute-hungry by nature, and local inference keeps the iteration loop fast and free.&lt;/p&gt;
&lt;hr /&gt;
&lt;h2&gt;What Can Evolve&lt;/h2&gt;
&lt;p&gt;Instructions are the most common target, but they're not the only evolvable component. gepa-adk can also optimize:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Output schemas&lt;/strong&gt; — the Pydantic model that structures the agent's response&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Generation config&lt;/strong&gt; — LLM parameters like temperature and top-p&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Multi-agent systems&lt;/strong&gt; — evolve instructions across multiple agents simultaneously, optimizing how they work together&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;That last one is where things get interesting. In a multi-agent pipeline, the instructions of one agent affect the inputs to the next. Evolving them together means the optimizer can find coordination patterns that no amount of individual prompt tuning would discover — like a kitchen brigade where the saucier and the grill cook learn to time their dishes together instead of each optimizing in isolation.&lt;/p&gt;
&lt;hr /&gt;
&lt;h2&gt;When to Use This&lt;/h2&gt;
&lt;p&gt;Evolutionary optimization isn't for every prompt. It shines when:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;You have measurable quality criteria&lt;/strong&gt; — if you can define what &amp;quot;good&amp;quot; looks like in a critic, evolution can optimize for it&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Your agent runs on diverse inputs&lt;/strong&gt; — evolution generalizes across training examples instead of overfitting to one&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Manual tuning has plateaued&lt;/strong&gt; — you've tweaked the prompt as far as your intuition goes&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;You're building for production&lt;/strong&gt; — the difference between a 0.65 and a 0.82 score matters when the agent runs thousands of times&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;It's less useful for one-off prompts, creative tasks without clear quality metrics, or situations where the &amp;quot;right answer&amp;quot; changes too frequently for a training set to be meaningful.&lt;/p&gt;
&lt;hr /&gt;
&lt;h2&gt;Key Takeaways&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Manual prompt engineering has a ceiling&lt;/strong&gt; — human intuition can't efficiently search the space of possible instructions&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Evolutionary optimization uses LLM-powered reflection&lt;/strong&gt; — not random mutation, but informed analysis of what's working and what isn't&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;A critic agent is the key ingredient&lt;/strong&gt; — structured scoring and feedback drive meaningful improvements&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;The evolved prompts are often surprising&lt;/strong&gt; — the algorithm finds instruction patterns you wouldn't write yourself&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Get started now&lt;/strong&gt;: &lt;code&gt;pip install gepa-adk&lt;/code&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;hr /&gt;
&lt;h2&gt;Further Reading&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href="https://github.com/Alberto-Codes/gepa-adk"&gt;gepa-adk on GitHub&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://pypi.org/project/gepa-adk/"&gt;gepa-adk on PyPI&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://alberto-codes.github.io/gepa-adk/"&gt;gepa-adk Documentation&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://arxiv.org/abs/2507.19457"&gt;GEPA Paper&lt;/a&gt; — the research behind the algorithm&lt;/li&gt;
&lt;li&gt;&lt;a href="https://google.github.io/adk-docs/"&gt;Google Agent Development Kit&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
</content:encoded>
    </item>
    <item>
      <title>Give Your AI Agent a Docstring Quality Tool</title>
      <link>https://alberto.codes/blog/2026-03-04-give-your-ai-agent-a-docstring-quality-tool</link>
      <guid isPermaLink="true">https://alberto.codes/blog/2026-03-04-give-your-ai-agent-a-docstring-quality-tool</guid>
      <pubDate>Wed, 04 Mar 2026 00:00:00 GMT</pubDate>
      <description>Wire docvet's MCP server into VS Code, Cursor, or Claude Code — your AI coding agent gets structured docstring quality checks without parsing CLI output.</description>
      <content:encoded>&lt;p&gt;Your AI coding agent can read your code, run your tests, and search your repo. But can it check whether your docstrings actually match what the code does?&lt;/p&gt;
&lt;p&gt;&lt;a href="https://alberto.codes/blog/2026-02-25-your-ai-reads-your-docstrings"&gt;In a previous post&lt;/a&gt;, I made the case that stale docstrings actively mislead AI agents — in controlled tests, incorrect documentation drops LLM task success by 22.6 percentage points. Agents given wrong docstrings performed worse than agents given no docstrings at all. &lt;a href="https://github.com/Alberto-Codes/docvet"&gt;docvet&lt;/a&gt; was built to catch those gaps: 19 rules that check whether your docstrings actually match what the code does.&lt;/p&gt;
&lt;p&gt;Since v1.8, docvet ships an MCP server. That means any MCP-aware editor — VS Code, Cursor, Claude Code, Windsurf, Claude Desktop — can give its AI agent direct access to docstring quality checks. Structured JSON, not CLI text parsing. One config block, zero install.&lt;/p&gt;
&lt;hr /&gt;
&lt;h2&gt;What Your Agent Gets&lt;/h2&gt;
&lt;p&gt;Two tools appear in the agent's toolbox:&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;&lt;code&gt;docvet_check&lt;/code&gt;&lt;/strong&gt; — Run checks on any Python file or directory. The agent receives structured findings:&lt;/p&gt;
&lt;pre&gt;&lt;code class="language-json"&gt;{
  &amp;quot;findings&amp;quot;: [
    {
      &amp;quot;file&amp;quot;: &amp;quot;src/pipeline/extract.py&amp;quot;,
      &amp;quot;line&amp;quot;: 42,
      &amp;quot;symbol&amp;quot;: &amp;quot;extract_text&amp;quot;,
      &amp;quot;rule&amp;quot;: &amp;quot;missing-raises&amp;quot;,
      &amp;quot;message&amp;quot;: &amp;quot;Function 'extract_text' raises ValueError but has no Raises section&amp;quot;,
      &amp;quot;category&amp;quot;: &amp;quot;required&amp;quot;
    }
  ],
  &amp;quot;summary&amp;quot;: {
    &amp;quot;total&amp;quot;: 3,
    &amp;quot;by_category&amp;quot;: {&amp;quot;required&amp;quot;: 2, &amp;quot;recommended&amp;quot;: 1},
    &amp;quot;files_checked&amp;quot;: 8
  }
}
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;&lt;strong&gt;&lt;code&gt;docvet_rules&lt;/code&gt;&lt;/strong&gt; — List all 19 rules with descriptions and categories. Useful when the agent needs to explain a finding or decide what to fix first.&lt;/p&gt;
&lt;p&gt;No CLI output to parse. No regex. The agent gets typed fields it can reason about directly.&lt;/p&gt;
&lt;hr /&gt;
&lt;h2&gt;Why MCP, Not CLI?&lt;/h2&gt;
&lt;p&gt;Without MCP, your agent shells out to &lt;code&gt;docvet check&lt;/code&gt; and has to make sense of this:&lt;/p&gt;
&lt;pre&gt;&lt;code class="language-console"&gt;src/pipeline/extract.py:42: missing-raises - Function 'extract_text' raises ValueError but has no Raises section
src/pipeline/extract.py:87: missing-param - Parameter 'timeout' not documented
src/utils/cache.py:15: stale-signature - Docstring parameters don't match function signature
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;That works, but the agent is regex-matching strings instead of reasoning about typed data. With MCP, it gets the structured JSON shown above — fields it can filter, sort, and act on directly.&lt;/p&gt;
&lt;hr /&gt;
&lt;h2&gt;Setup: One Block of JSON&lt;/h2&gt;
&lt;p&gt;The MCP server runs on stdio via &lt;code&gt;uvx&lt;/code&gt; — no &lt;code&gt;pip install&lt;/code&gt; in your project, no virtual environment pollution, no global packages. &lt;code&gt;uvx&lt;/code&gt; downloads and runs docvet in an isolated environment automatically. You add the config and it just works.&lt;/p&gt;
&lt;h3&gt;VS Code&lt;/h3&gt;
&lt;p&gt;Add to &lt;code&gt;.vscode/mcp.json&lt;/code&gt; in your project root:&lt;/p&gt;
&lt;pre&gt;&lt;code class="language-json"&gt;{
  &amp;quot;servers&amp;quot;: {
    &amp;quot;docvet&amp;quot;: {
      &amp;quot;type&amp;quot;: &amp;quot;stdio&amp;quot;,
      &amp;quot;command&amp;quot;: &amp;quot;uvx&amp;quot;,
      &amp;quot;args&amp;quot;: [&amp;quot;docvet[mcp]&amp;quot;, &amp;quot;mcp&amp;quot;]
    }
  }
}
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Note: VS Code uses &lt;code&gt;&amp;quot;servers&amp;quot;&lt;/code&gt; as the top-level key, not &lt;code&gt;&amp;quot;mcpServers&amp;quot;&lt;/code&gt;.&lt;/p&gt;
&lt;h3&gt;Cursor&lt;/h3&gt;
&lt;p&gt;Add to &lt;code&gt;.cursor/mcp.json&lt;/code&gt; (project-level) or &lt;code&gt;~/.cursor/mcp.json&lt;/code&gt; (global):&lt;/p&gt;
&lt;pre&gt;&lt;code class="language-json"&gt;{
  &amp;quot;mcpServers&amp;quot;: {
    &amp;quot;docvet&amp;quot;: {
      &amp;quot;command&amp;quot;: &amp;quot;uvx&amp;quot;,
      &amp;quot;args&amp;quot;: [&amp;quot;docvet[mcp]&amp;quot;, &amp;quot;mcp&amp;quot;]
    }
  }
}
&lt;/code&gt;&lt;/pre&gt;
&lt;h3&gt;Claude Code&lt;/h3&gt;
&lt;p&gt;One command, no JSON editing:&lt;/p&gt;
&lt;pre&gt;&lt;code class="language-bash"&gt;claude mcp add --transport stdio --scope project docvet -- uvx &amp;quot;docvet[mcp]&amp;quot; mcp
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Use &lt;code&gt;--scope user&lt;/code&gt; to make it available across all projects.&lt;/p&gt;
&lt;h3&gt;Windsurf / Claude Desktop&lt;/h3&gt;
&lt;p&gt;Same &lt;code&gt;mcpServers&lt;/code&gt; pattern — see the &lt;a href="https://alberto-codes.github.io/docvet/editor-integration/#client-configuration"&gt;full editor guide&lt;/a&gt; for copy-paste configs. There's also a &lt;a href="https://alberto-codes.github.io/docvet/editor-integration/#__tabbed_1_3"&gt;dedicated VS Code setup tab&lt;/a&gt; if you just want the one-click config.&lt;/p&gt;
&lt;hr /&gt;
&lt;h2&gt;What Happens Next&lt;/h2&gt;
&lt;p&gt;Once configured, your AI agent can invoke &lt;code&gt;docvet_check&lt;/code&gt; on any file it's working with. A typical workflow:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;Agent opens a Python file to modify&lt;/li&gt;
&lt;li&gt;Agent runs &lt;code&gt;docvet_check&lt;/code&gt; on the file&lt;/li&gt;
&lt;li&gt;Findings come back as structured JSON — missing Raises sections, stale signatures, undocumented attributes&lt;/li&gt;
&lt;li&gt;Agent fixes the docstrings alongside the code change&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;The feedback loop becomes automatic — like a line cook who taste-tests every dish before it leaves the pass. The agent checks docstring quality as part of its normal workflow, not as an afterthought.&lt;/p&gt;
&lt;hr /&gt;
&lt;h2&gt;Try It&lt;/h2&gt;
&lt;p&gt;If you have VS Code with MCP support enabled:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;Add the &lt;code&gt;.vscode/mcp.json&lt;/code&gt; block above&lt;/li&gt;
&lt;li&gt;Open a Python file with a known docstring gap (a function that raises an exception but has no &lt;code&gt;Raises:&lt;/code&gt; section)&lt;/li&gt;
&lt;li&gt;Ask your AI agent to check the file with docvet&lt;/li&gt;
&lt;li&gt;Watch it get structured findings and fix the docstring&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;Now your agent can check whether your docstrings actually match what the code does. One config block, zero install.&lt;/p&gt;
&lt;p&gt;docvet is on &lt;a href="https://pypi.org/project/docvet/"&gt;PyPI&lt;/a&gt;, &lt;a href="https://github.com/Alberto-Codes/docvet"&gt;GitHub&lt;/a&gt;, and the &lt;a href="https://registry.modelcontextprotocol.io"&gt;MCP Registry&lt;/a&gt;. The full editor setup guide is at &lt;a href="https://alberto-codes.github.io/docvet/editor-integration/"&gt;alberto-codes.github.io/docvet/editor-integration&lt;/a&gt;.&lt;/p&gt;
&lt;hr /&gt;
&lt;h2&gt;Further Reading&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href="https://alberto-codes.github.io/docvet/editor-integration/#__tabbed_1_3"&gt;VS Code MCP setup (copy-paste config)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://alberto-codes.github.io/docvet/editor-integration/"&gt;Full editor integration docs&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/Alberto-Codes/docvet/discussions"&gt;GitHub announcement: docvet is on the MCP Registry&lt;/a&gt; — v1.8 launch post&lt;/li&gt;
&lt;li&gt;&lt;a href="https://alberto.codes/blog/2026-02-25-your-ai-reads-your-docstrings"&gt;Your AI Reads Your Docstrings. Are They Right?&lt;/a&gt; — the problem this solves&lt;/li&gt;
&lt;li&gt;&lt;a href="https://modelcontextprotocol.io/"&gt;MCP specification&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://code.visualstudio.com/docs/copilot/chat/mcp-servers"&gt;VS Code MCP support&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
</content:encoded>
    </item>
    <item>
      <title>Your AI Agent's Memories Aren't Encrypted</title>
      <link>https://alberto.codes/blog/2026-03-01-your-ai-agents-memories-arent-encrypted</link>
      <guid isPermaLink="true">https://alberto.codes/blog/2026-03-01-your-ai-agents-memories-arent-encrypted</guid>
      <pubDate>Sun, 01 Mar 2026 00:00:00 GMT</pubDate>
      <description>Google ADK stores everything your agent knows — tool calls, user messages, conversation context — in plaintext SQLite. Here's why that matters and how to fix it.</description>
      <content:encoded>&lt;p&gt;You built an AI agent. It calls tools, remembers conversations, and tracks state across sessions. It works. You deploy it. Users start talking to it — sharing API keys, asking about internal systems, passing along customer data through tool calls.&lt;/p&gt;
&lt;p&gt;Every word of that is sitting in a SQLite file, in plaintext, readable by anyone with file access.&lt;/p&gt;
&lt;p&gt;This isn't a vulnerability you introduced. It's the default behavior of Google ADK's session persistence. And if you're building agents that handle anything beyond demo data, it's a problem you need to solve before production.&lt;/p&gt;
&lt;hr /&gt;
&lt;h2&gt;The Problem You Don't Know You Have&lt;/h2&gt;
&lt;p&gt;Google's Agent Development Kit gives you &lt;code&gt;DatabaseSessionService&lt;/code&gt; — a session persistence layer that stores agent state in a relational database. It handles the hard parts: session lifecycle, event ordering, state reconstruction across turns. You configure it with a connection string and move on to the interesting work of building your agent.&lt;/p&gt;
&lt;pre&gt;&lt;code class="language-python"&gt;from google.adk.sessions import DatabaseSessionService

session_service = DatabaseSessionService(db_url=&amp;quot;sqlite+aiosqlite:///sessions.db&amp;quot;)
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;That &lt;code&gt;sessions.db&lt;/code&gt; file now contains everything your agent processes. Every user message. Every tool call and its arguments. Every function response. Every piece of state your agent carries between turns. All of it stored as structured data, fully readable, with no encryption layer between the file system and your users' data.&lt;/p&gt;
&lt;p&gt;Open the file. Run a query. Read the JSON. There's nothing stopping you — and nothing stopping anyone else who gains access to that file.&lt;/p&gt;
&lt;hr /&gt;
&lt;h2&gt;What's Actually in a Session?&lt;/h2&gt;
&lt;p&gt;ADK sessions aren't just chat logs. They're structured records of everything your agent does:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;User messages&lt;/strong&gt; — the full text of every prompt your users send&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Tool calls&lt;/strong&gt; — function names and their complete argument payloads&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Function responses&lt;/strong&gt; — the raw data your tools return, including API responses and database query results&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Agent state&lt;/strong&gt; — accumulated context the agent carries across turns&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Event metadata&lt;/strong&gt; — timestamps, ordering, session identifiers&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;img src="https://alberto.codes/adk-session-data-flow.svg" alt="Data flows through an ADK agent session — from user messages and tool calls to storage. DatabaseSessionService stores it as readable JSON. EncryptedSessionService stores it as encrypted blobs." /&gt;&lt;/p&gt;
&lt;p&gt;Think about what flows through a real agent. A user asks your agent to look up a customer record — the tool call contains the query, and the response contains the customer's data. A user passes an API key so the agent can call an external service — that key is now stored in the session. A user discusses internal business logic — the agent remembers it for context, and the database remembers it forever.&lt;/p&gt;
&lt;p&gt;And none of this is hypothetical. If you open that SQLite file with any database browser, every session row contains JSON you can read without any special tooling. No keys, no decryption step, no access control. The file is the data, and the data is in the clear.&lt;/p&gt;
&lt;p&gt;The session database isn't a chat transcript. It's a complete operational record of everything your agent knows about your users and what it did on their behalf. It's like a kitchen where every order ticket, every ingredient substitution, and every customer allergy note gets pinned to the wall in plain sight — and the back door is unlocked.&lt;/p&gt;
&lt;hr /&gt;
&lt;h2&gt;Why This Matters Now&lt;/h2&gt;
&lt;p&gt;When agents live in notebooks and demos, plaintext sessions are fine. Nobody's passing real credentials to a prototype. But agents are leaving the sandbox. It's the difference between cooking at home and running a restaurant — in your own kitchen, nobody cares if the pantry is unlocked. The moment you're serving customers, health inspectors show up and every storage decision matters. Google released ADK to help developers build production-grade agents, and teams are taking that seriously — building agents that handle real workflows with real data.&lt;/p&gt;
&lt;p&gt;Production agents process real user data. They call real APIs with real keys. They handle internal workflows where the conversation itself is sensitive. The moment your agent crosses from &amp;quot;demo&amp;quot; to &amp;quot;deployed,&amp;quot; the session database becomes a data-at-rest liability.&lt;/p&gt;
&lt;p&gt;Compliance frameworks are explicit about this:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;SOC 2&lt;/strong&gt; requires encryption of sensitive data at rest&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;HIPAA&lt;/strong&gt; mandates encryption for protected health information in storage&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;GDPR&lt;/strong&gt; expects appropriate technical measures to protect personal data&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;These aren't edge cases. If your agent processes user data and you're pursuing any compliance certification, an auditor will ask how session data is protected at rest. &amp;quot;It's in a SQLite file&amp;quot; is not an answer that passes review.&lt;/p&gt;
&lt;p&gt;&amp;quot;The agent works&amp;quot; is necessary. &amp;quot;The agent is secure&amp;quot; is what lets you ship it. And right now, ADK gives you persistence but not protection. That gap between working and secure is exactly where production agents get stuck.&lt;/p&gt;
&lt;hr /&gt;
&lt;h2&gt;The Fix Is One Import Away&lt;/h2&gt;
&lt;p&gt;&lt;a href="https://github.com/Alberto-Codes/adk-secure-sessions"&gt;adk-secure-sessions&lt;/a&gt; is an encrypted session storage layer for Google ADK. It's a drop-in replacement — same session interface, same behavior, encrypted storage.&lt;/p&gt;
&lt;pre&gt;&lt;code class="language-python"&gt;from adk_secure_sessions import EncryptedSessionService, FernetBackend

session_service = EncryptedSessionService(
    db_url=&amp;quot;sqlite+aiosqlite:///sessions.db&amp;quot;,
    backend=FernetBackend(&amp;quot;your-secret-key&amp;quot;),
)
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Your agent code doesn't change. Every call to create, get, list, or delete sessions works the same way. The difference is what happens at the storage layer: session data is serialized, encrypted, then written. On read, it's decrypted and deserialized. The SQLite file contains encrypted blobs instead of readable JSON.&lt;/p&gt;
&lt;p&gt;The encryption is Fernet — AES-128-CBC with HMAC-SHA256 for authenticated encryption, meaning data is both confidential and tamper-evident. Key derivation uses PBKDF2-HMAC-SHA256 with 480,000 iterations, so your passphrase becomes a strong cryptographic key without managing raw key material.&lt;/p&gt;
&lt;p&gt;Two runtime dependencies: &lt;code&gt;google-adk&lt;/code&gt; and &lt;code&gt;cryptography&lt;/code&gt;. Nothing exotic. Nothing heavy. No C extensions to compile, no system libraries to install.&lt;/p&gt;
&lt;p&gt;The library extends ADK's &lt;code&gt;BaseSessionService&lt;/code&gt; directly rather than wrapping &lt;code&gt;DatabaseSessionService&lt;/code&gt;. That's a deliberate architectural choice — and one that came from a failed first attempt with the decorator pattern. That story, and the generalizable lesson about when composition breaks down, are coming in Part 4 of this series.&lt;/p&gt;
&lt;hr /&gt;
&lt;h2&gt;What Changes, What Doesn't&lt;/h2&gt;
&lt;p&gt;What stays the same:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Your agent code — no changes to how you interact with sessions&lt;/li&gt;
&lt;li&gt;Session behavior — create, read, update, delete all work identically&lt;/li&gt;
&lt;li&gt;The database — still SQLite, still a single file, still portable&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;What changes:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Session data at rest is encrypted — the file is useless without the key&lt;/li&gt;
&lt;li&gt;Error messages are precise — &lt;code&gt;DecryptionError&lt;/code&gt; and &lt;code&gt;ConfigurationError&lt;/code&gt; tell you exactly what went wrong&lt;/li&gt;
&lt;li&gt;You need a passphrase — managed via environment variable, not hardcoded&lt;/li&gt;
&lt;/ul&gt;
&lt;pre&gt;&lt;code class="language-python"&gt;import os
from adk_secure_sessions import EncryptedSessionService, FernetBackend

session_service = EncryptedSessionService(
    db_url=&amp;quot;sqlite+aiosqlite:///sessions.db&amp;quot;,
    backend=FernetBackend(os.environ[&amp;quot;SESSION_KEY&amp;quot;]),
)
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;That's the production pattern. The passphrase lives in your environment, the encrypted data lives in your database, and the two never meet in your codebase. Standard secrets management — nothing new to learn, nothing unusual to maintain.&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;With &lt;code&gt;DatabaseSessionService&lt;/code&gt;, an attacker with file access sees everything. With &lt;code&gt;EncryptedSessionService&lt;/code&gt;, they see noise.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;hr /&gt;
&lt;h2&gt;Key Takeaways&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Google ADK stores session state in plaintext by default&lt;/strong&gt; — every tool call, user message, and piece of agent context is readable in the database&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Sessions contain more than chat history&lt;/strong&gt; — tool arguments, API responses, and accumulated state are all persisted&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Production agents need encryption at rest&lt;/strong&gt; — compliance frameworks require it, and good engineering practice demands it&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;adk-secure-sessions is a drop-in replacement&lt;/strong&gt; — one import change, same API, encrypted storage&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Get started now&lt;/strong&gt;: &lt;code&gt;pip install adk-secure-sessions&lt;/code&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;hr /&gt;
&lt;h2&gt;Further Reading&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href="https://github.com/Alberto-Codes/adk-secure-sessions"&gt;adk-secure-sessions on GitHub&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://pypi.org/project/adk-secure-sessions/"&gt;adk-secure-sessions on PyPI&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://google.github.io/adk-docs/"&gt;Google Agent Development Kit&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;Next in this series: Encrypting ADK Sessions in 5 Minutes (Part 2)&lt;/li&gt;
&lt;/ul&gt;
</content:encoded>
    </item>
    <item>
      <title>Your AI Reads Your Docstrings. Are They Right?</title>
      <link>https://alberto.codes/blog/2026-02-25-your-ai-reads-your-docstrings</link>
      <guid isPermaLink="true">https://alberto.codes/blog/2026-02-25-your-ai-reads-your-docstrings</guid>
      <pubDate>Wed, 25 Feb 2026 00:00:00 GMT</pubDate>
      <description>Stale docstrings poison your AI coding agent's understanding of your codebase. Research shows incorrect documentation is worse than no documentation at all.</description>
      <content:encoded>&lt;p&gt;You're pair-programming with an AI agent. It reads your codebase, finds the function you're extending, checks the docstring, and generates code that calls it with the old parameter names. You debug for twenty minutes before realizing the docstring was updated six months ago—wait, no. The &lt;em&gt;function&lt;/em&gt; was updated six months ago. The docstring still describes the version before the refactor.&lt;/p&gt;
&lt;p&gt;This is the software equivalent of a recipe card that says &amp;quot;bake at 350°F for 30 minutes&amp;quot; when someone already modified the dish to be pan-fried. A cook following that card doesn't just get a mediocre result—they ruin the dish entirely. Outdated instructions are worse than no instructions, because they create false confidence.&lt;/p&gt;
&lt;p&gt;Your AI coding agent reads your docstrings. Every time it generates code, it uses those docstrings as context for understanding your functions, classes, and modules. When those docstrings are stale, incomplete, or wrong, the agent doesn't know to distrust them. It follows the recipe card.&lt;/p&gt;
&lt;hr /&gt;
&lt;h2&gt;The Research Confirms What You Already Suspect&lt;/h2&gt;
&lt;p&gt;This isn't just an intuition. Multiple studies have quantified how documentation quality affects LLM performance on code tasks:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Incorrect documentation degrades LLM task success by 22.6 percentage points&lt;/strong&gt; compared to correct docs (&lt;a href="https://arxiv.org/abs/2404.03114"&gt;Macke &amp;amp; Doyle, 2024&lt;/a&gt;). Critically, incomplete or missing docs didn't cause nearly the same harm—wrong docs are uniquely toxic.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Comment density improves code generation by 40–54%&lt;/strong&gt; across benchmarks (&lt;a href="https://arxiv.org/abs/2402.13013"&gt;Wei et al., 2024&lt;/a&gt;). More documentation means better AI output, but only when that documentation is accurate.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Misleading comments reduce LLM fault localization accuracy to 24.55%&lt;/strong&gt; (&lt;a href="https://arxiv.org/abs/2504.04372"&gt;Jia et al., 2025&lt;/a&gt;). When the comments lie, the AI can't find bugs.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Performance drops substantially without docstrings&lt;/strong&gt;, and intent-aware inference is needed to compensate (&lt;a href="https://arxiv.org/abs/2508.09537"&gt;Li et al., 2025&lt;/a&gt;).&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;The 2025 DORA report puts it bluntly: &lt;em&gt;&amp;quot;AI doesn't fix a team; it amplifies what's already there.&amp;quot;&lt;/em&gt; If your docstrings are accurate, AI makes you faster. If they're stale, AI confidently generates the wrong code—and you trust it because the agent seemed so sure.&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Stale docstrings don't just fail to help. They actively mislead. Incorrect documentation is worse than no documentation at all.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;Think of it like a pantry with mislabeled jars. A cook with an empty shelf knows to go find ingredients. A cook with a jar labeled &amp;quot;cumin&amp;quot; that's actually filled with cinnamon? That cook seasons the chili, tastes nothing wrong until it's too late, and serves a dish that's subtly, confusingly off. That's what stale docstrings do to your AI agent.&lt;/p&gt;
&lt;hr /&gt;
&lt;h2&gt;The Gap in Your Toolchain&lt;/h2&gt;
&lt;p&gt;You probably already have docstring tooling. Most Python projects use some combination of:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;ruff&lt;/strong&gt; (D rules) — checks how your docstrings &lt;em&gt;look&lt;/em&gt;. Formatting, style conventions, section order.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;interrogate&lt;/strong&gt; — checks if docstrings &lt;em&gt;exist&lt;/em&gt;. Coverage percentage across your codebase.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;These are layers 1 and 2 of docstring quality. Necessary, but not sufficient. They answer &amp;quot;is there a docstring?&amp;quot; and &amp;quot;is it formatted correctly?&amp;quot; They don't answer the question that actually matters for AI agents:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Is the docstring &lt;em&gt;right&lt;/em&gt;?&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;Docstring quality has six distinct layers:&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Layer&lt;/th&gt;
&lt;th&gt;Question&lt;/th&gt;
&lt;th&gt;Tool&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;1. Presence&lt;/td&gt;
&lt;td&gt;Does it exist?&lt;/td&gt;
&lt;td&gt;interrogate&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2. Style&lt;/td&gt;
&lt;td&gt;Is it formatted correctly?&lt;/td&gt;
&lt;td&gt;ruff D rules&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;3. Completeness&lt;/td&gt;
&lt;td&gt;Does it document all sections?&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;gap&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;4. Accuracy&lt;/td&gt;
&lt;td&gt;Does it match the current code?&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;gap&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;5. Rendering&lt;/td&gt;
&lt;td&gt;Will mkdocs render it correctly?&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;gap&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;6. Visibility&lt;/td&gt;
&lt;td&gt;Will mkdocs even see the file?&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;gap&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;&lt;img src="https://alberto.codes/docstring-quality-layers.svg" alt="Six layers of docstring quality: presence and style are covered by existing tools, while completeness, accuracy, rendering, and visibility remain unchecked" /&gt;&lt;/p&gt;
&lt;p&gt;Layers 1–2 are table stakes. Layers 3–6 are where your AI agent's understanding lives or dies—and until now, no tool covered them.&lt;/p&gt;
&lt;p&gt;It's like a restaurant kitchen where health inspectors check that you &lt;em&gt;have&lt;/em&gt; a recipe binder (presence) and that the cards are &lt;em&gt;legible&lt;/em&gt; (style), but nobody ever verifies that the recipes match what the cooks are actually preparing. The card says &amp;quot;sear for 2 minutes per side&amp;quot; but the chef switched to a 4-minute sear last month. The binder looks great. The food is wrong.&lt;/p&gt;
&lt;hr /&gt;
&lt;h2&gt;Enter docvet&lt;/h2&gt;
&lt;p&gt;&lt;a href="https://github.com/Alberto-Codes/docvet"&gt;docvet&lt;/a&gt; fills layers 3–6 with 19 rules across four checks:&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Enrichment&lt;/strong&gt; (10 rules) — completeness. Your function raises &lt;code&gt;ValueError&lt;/code&gt; but the docstring has no &lt;code&gt;Raises:&lt;/code&gt; section? Your dataclass has five attributes but no &lt;code&gt;Attributes:&lt;/code&gt; section? docvet catches it. It reads your AST—the actual code structure—and compares it against what the docstring claims to document.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Freshness&lt;/strong&gt; (5 rules) — accuracy. This is the killer feature. docvet uses &lt;code&gt;git diff&lt;/code&gt; and &lt;code&gt;git blame&lt;/code&gt; to detect when code changes but docstrings don't. Changed a function's signature last week? The docstring still describes the old parameters? That's a &lt;code&gt;stale-signature&lt;/code&gt; finding, severity HIGH.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Griffe&lt;/strong&gt; (3 rules) — rendering compatibility. If you publish docs with mkdocs, docvet catches griffe parser warnings before they silently break your documentation site.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Coverage&lt;/strong&gt; (1 rule) — visibility. Missing &lt;code&gt;__init__.py&lt;/code&gt; files make entire packages invisible to documentation generators. docvet finds them.&lt;/p&gt;
&lt;p&gt;One line to try it:&lt;/p&gt;
&lt;pre&gt;&lt;code class="language-bash"&gt;pip install docvet &amp;amp;&amp;amp; docvet check --all
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Example output on a real codebase:&lt;/p&gt;
&lt;pre&gt;&lt;code class="language-text"&gt;src/pipeline/extract.py:42: stale-signature Function 'extract_text' signature changed but docstring not updated [required]
src/models/customer.py:15: missing-attributes Dataclass 'CustomerRecord' has no Attributes: section [required]
src/utils/validate.py:88: missing-raises Function 'validate_schema' raises ValueError but has no Raises: section [required]

3 findings (3 required, 0 recommended)
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Each finding tells you exactly what's wrong, where it is, and whether it's required or recommended. No configuration needed to start—docvet runs with sensible defaults out of the box.&lt;/p&gt;
&lt;hr /&gt;
&lt;h2&gt;The Feedback Loop&lt;/h2&gt;
&lt;p&gt;The tagline &amp;quot;better docstrings, better AI&amp;quot; isn't marketing—it's a literal feedback loop:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;AI reads your docstrings to understand your code&lt;/li&gt;
&lt;li&gt;docvet ensures those docstrings are complete and accurate&lt;/li&gt;
&lt;li&gt;AI writes better code because its context is trustworthy&lt;/li&gt;
&lt;li&gt;Repeat&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;If your AI agent is your sous chef, your docstrings are the recipe cards pinned above each station. docvet is the head chef who walks the line before service, pulls down every card, and checks it against what's actually in the pan. No stale cards make it to service.&lt;/p&gt;
&lt;hr /&gt;
&lt;h2&gt;Key Takeaways&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Your AI coding agent reads your docstrings&lt;/strong&gt; as primary context for understanding your codebase—stale docs mean bad AI output&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Incorrect documentation is worse than no documentation&lt;/strong&gt;—research shows a 22.6 percentage point drop in LLM task success with wrong docs&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Existing tools cover style and presence&lt;/strong&gt; (layers 1–2) but not completeness, accuracy, rendering, or visibility (layers 3–6)&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;docvet fills the gap&lt;/strong&gt; with 19 rules across four checks: enrichment, freshness, griffe, and coverage&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Freshness detection uses git history&lt;/strong&gt; to catch code-docstring drift—no manual review needed&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;One command to try it&lt;/strong&gt;: &lt;code&gt;pip install docvet &amp;amp;&amp;amp; docvet check --all&lt;/code&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;hr /&gt;
&lt;h2&gt;Further Reading&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href="https://github.com/Alberto-Codes/docvet"&gt;docvet on GitHub&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://pypi.org/project/docvet/"&gt;docvet on PyPI&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://alberto-codes.github.io/docvet/"&gt;docvet Documentation&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://arxiv.org/abs/2404.03114"&gt;Testing the Effect of Code Documentation on LLM Code Understanding (Macke &amp;amp; Doyle, 2024)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://arxiv.org/abs/2402.13013"&gt;Code Needs Comments: Enhancing Code LLMs with Comment Augmentation (Wei et al., 2024)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://arxiv.org/abs/2504.04372"&gt;Assessing the Impact of Code Changes on Fault Localizability of LLMs (Jia et al., 2025)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://arxiv.org/abs/2508.09537"&gt;Your Coding Intent is Secretly in the Context (Li et al., 2025)&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
</content:encoded>
    </item>
    <item>
      <title>Entity Resolution is Recipe Matching</title>
      <link>https://alberto.codes/blog/2026-02-05-entity-resolution-is-recipe-matching</link>
      <guid isPermaLink="true">https://alberto.codes/blog/2026-02-05-entity-resolution-is-recipe-matching</guid>
      <pubDate>Thu, 05 Feb 2026 00:00:00 GMT</pubDate>
      <description>When exact matching fails, probabilistic record linkage weighs evidence like a chef recognizes a dish—not by a single ingredient, but by the whole picture.</description>
      <content:encoded>&lt;p&gt;Birria in Jalisco is goat, slow-braised in dried chiles and served in a bowl with consommé. In Tijuana, it's beef, crisped on a plancha and folded into a taco. At Chipotle, it's a menu item with queso. Three different presentations, three different proteins, three different contexts—but you recognize them as the same dish.&lt;/p&gt;
&lt;p&gt;You're not matching on a single ingredient. You're weighing evidence across multiple dimensions: the chile base, the braising technique, the way it's served with its cooking liquid.&lt;/p&gt;
&lt;p&gt;Entity resolution works the same way. When you have two database records—&amp;quot;José García, 123 Main St, DOB 1985-03-15&amp;quot; and &amp;quot;GARCIA, JOSE M., 123 Main Street, DOB 03/15/1985&amp;quot;—you need to determine if they represent the same person. No single field gives you certainty, but together, the evidence points strongly toward a match.&lt;/p&gt;
&lt;p&gt;I solved this problem on a project in financial services: deduplicating customer records, linking transactions across systems, matching entities without shared keys. It's harder than it looks, and most engineers reach for exact matching first—which works until it doesn't.&lt;/p&gt;
&lt;hr /&gt;
&lt;h2&gt;Why Exact Matching Fails&lt;/h2&gt;
&lt;p&gt;The instinct is to write rules:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;If name matches &lt;strong&gt;AND&lt;/strong&gt; DOB matches &lt;strong&gt;AND&lt;/strong&gt; address matches → same person&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;This is &lt;strong&gt;deterministic matching&lt;/strong&gt;, and it breaks the moment it touches real-world data.&lt;/p&gt;
&lt;p&gt;Consider what customer data actually looks like:&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Field&lt;/th&gt;
&lt;th&gt;System A&lt;/th&gt;
&lt;th&gt;System B&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Name&lt;/td&gt;
&lt;td&gt;José García&lt;/td&gt;
&lt;td&gt;GARCIA, JOSE M.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Address&lt;/td&gt;
&lt;td&gt;123 Main St&lt;/td&gt;
&lt;td&gt;123 Main Street, Apt 2B&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;DOB&lt;/td&gt;
&lt;td&gt;1985-03-15&lt;/td&gt;
&lt;td&gt;03/15/1985&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Email&lt;/td&gt;
&lt;td&gt;jose.garcia@gmail.com&lt;/td&gt;
&lt;td&gt;jgarcia85@yahoo.com&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;Exact matching finds &lt;strong&gt;zero matches&lt;/strong&gt; here. Every field differs in format, abbreviation, or completeness. But a human reviewer immediately recognizes this as likely the same person.&lt;/p&gt;
&lt;p&gt;You could write transformation rules—normalize names to uppercase, expand abbreviations, standardize date formats. This helps, but it's a game of whack-a-mole:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&amp;quot;José&amp;quot; vs &amp;quot;Jose&amp;quot; vs &amp;quot;Joe&amp;quot;&lt;/li&gt;
&lt;li&gt;&amp;quot;García&amp;quot; vs &amp;quot;Garcia&amp;quot;&lt;/li&gt;
&lt;li&gt;&amp;quot;123 Main&amp;quot; vs &amp;quot;123 Main St&amp;quot; vs &amp;quot;123 Main Street&amp;quot;&lt;/li&gt;
&lt;li&gt;Typos: &amp;quot;Josè Garçia&amp;quot; → &amp;quot;José García&amp;quot;&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Deterministic matching forces a binary decision: match or no match. Reality is probabilistic—some pairs are definite matches, some definite non-matches, and a large gray zone requires weighing evidence.&lt;/p&gt;
&lt;p&gt;&lt;img src="https://alberto.codes/deterministic-vs-probabilistic.svg" alt="Deterministic matching gives binary yes/no decisions, while probabilistic matching produces weighted scores with room for manual review" /&gt;&lt;/p&gt;
&lt;hr /&gt;
&lt;h2&gt;Probabilistic Matching: Weighing Evidence&lt;/h2&gt;
&lt;p&gt;Probabilistic record linkage flips the question. Instead of asking &amp;quot;do these records match?&amp;quot; it asks:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;How likely is it that these records refer to the same entity, given the evidence?&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;Each field comparison contributes evidence:&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Comparison&lt;/th&gt;
&lt;th&gt;Evidence&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Exact name match&lt;/td&gt;
&lt;td&gt;Strong positive&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Fuzzy name match (Jaro-Winkler &amp;gt; 0.9)&lt;/td&gt;
&lt;td&gt;Moderate positive&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;DOB matches&lt;/td&gt;
&lt;td&gt;Strong positive&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;DOB off by one digit&lt;/td&gt;
&lt;td&gt;Weak positive (likely typo)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Completely different DOB&lt;/td&gt;
&lt;td&gt;Strong negative&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;The accumulated evidence produces a &lt;strong&gt;match score&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;This mirrors how you recognize birria. If two recipes both have dried chiles, braised meat, and a rich cooking liquid served alongside, that's strong evidence they're the same dish. If one uses guajillo and the other uses ancho, that's weak variation—regional preference, not a different dish. If one serves it in a bowl and the other crisps it in a taco, that's presentation, not identity.&lt;/p&gt;
&lt;p&gt;But if one is a quick stovetop stew without the slow braise? That's evidence you're looking at something else entirely—maybe carne guisada, but not birria.&lt;/p&gt;
&lt;hr /&gt;
&lt;h2&gt;The Fellegi-Sunter Model (Without the Math)&lt;/h2&gt;
&lt;p&gt;The math behind probabilistic matching is the &lt;strong&gt;Fellegi-Sunter model&lt;/strong&gt;, developed in 1969 and still the foundation of modern record linkage.&lt;/p&gt;
&lt;p&gt;Fellegi-Sunter assigns each field comparison a weight based on two probabilities:&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Probability&lt;/th&gt;
&lt;th&gt;Question&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;m-probability&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;If two records truly match, how often would this field comparison look like this?&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;u-probability&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;If two records &lt;em&gt;don't&lt;/em&gt; match, how often would this field comparison look like this by chance?&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;The ratio of these probabilities determines the weight:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Names match exactly&lt;/strong&gt;: Common among true matches (high m), rare among random pairs (low u) → &lt;strong&gt;positive weight&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Names completely differ&lt;/strong&gt;: Rare among true matches (low m), common among random pairs (high u) → &lt;strong&gt;negative weight&lt;/strong&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;h3&gt;Common vs. Rare Values&lt;/h3&gt;
&lt;p&gt;Here's the key insight: &lt;strong&gt;common values provide weaker evidence than rare values&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;&amp;quot;María García&amp;quot; matching &amp;quot;María García&amp;quot; is less compelling than &amp;quot;Xiomara Tlapoyawa&amp;quot; matching &amp;quot;Xiomara Tlapoyawa.&amp;quot; The first could happen by coincidence in any large dataset; the second almost certainly indicates the same person.&lt;/p&gt;
&lt;p&gt;Think of it like identifying a mole. If two recipes both contain chiles, that tells you almost nothing—every mole has chiles. But if both call for chocolate, banana, and chipotle in specific proportions, you're probably looking at variations of mole negro from Oaxaca.&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;The rare, specific ingredients carry more weight than the ubiquitous ones.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;Splink handles all of this automatically. You define what fields to compare and how (exact match, fuzzy match, within a date range), and Splink estimates the m and u probabilities using unsupervised learning—no labeled training data required.&lt;/p&gt;
&lt;hr /&gt;
&lt;h2&gt;The Blocking Problem&lt;/h2&gt;
&lt;p&gt;There's a computational catch. If you have a million records, comparing every pair means:&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;1,000,000 × 999,999 / 2 = 499,999,500,000 comparisons
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;That's &lt;strong&gt;500 billion comparisons&lt;/strong&gt;. Not feasible.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Blocking&lt;/strong&gt; solves this by only comparing records that could plausibly match. You define rules like:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Only compare records where first name initial AND birth year match&lt;/li&gt;
&lt;li&gt;Only compare records in the same city&lt;/li&gt;
&lt;li&gt;Only compare records with similar phone area codes&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;img src="https://alberto.codes/blocking-comparison.svg" alt="Without blocking: 500 billion comparisons taking weeks. With blocking: grouped comparisons completing in minutes" /&gt;&lt;/p&gt;
&lt;p&gt;This is &lt;strong&gt;mise en place&lt;/strong&gt; for data matching—organizing your workspace before you start cooking. A chef doesn't wander the entire kitchen looking for ingredients during service. They set up their station with everything they'll need in reach. Blocking sets up your comparison space so you're only looking where matches are likely to be found.&lt;/p&gt;
&lt;p&gt;The art is in the balance:&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Problem&lt;/th&gt;
&lt;th&gt;Cause&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Missing matches&lt;/td&gt;
&lt;td&gt;Blocking too aggressive—true pairs never compared&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Slow performance&lt;/td&gt;
&lt;td&gt;Blocking too loose—comparing pairs that obviously don't match&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;hr /&gt;
&lt;h2&gt;Why Splink?&lt;/h2&gt;
&lt;p&gt;You could implement Fellegi-Sunter from scratch, but &lt;a href="https://moj-analytical-services.github.io/splink/"&gt;Splink&lt;/a&gt; handles the hard parts:&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Feature&lt;/th&gt;
&lt;th&gt;What it does&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Unsupervised training&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Estimates m/u probabilities without labeled data (Expectation-Maximization)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Multiple backends&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;DuckDB (laptop), Spark (cluster), AWS Athena (serverless)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Fuzzy matching&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Jaro-Winkler, Levenshtein, phonetic matching built-in&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Term frequency&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Automatically down-weights common values like &amp;quot;José García&amp;quot;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Visualization&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Inspect match weights, debug blocking rules, understand model behavior&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;It's open source, built by the UK Ministry of Justice, and battle-tested on national-scale datasets including the 2021 Census.&lt;/p&gt;
&lt;hr /&gt;
&lt;h2&gt;Why This Matters in Financial Services&lt;/h2&gt;
&lt;p&gt;Entity resolution isn't an academic exercise. In financial services, getting it wrong has real consequences:&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Use Case&lt;/th&gt;
&lt;th&gt;Cost of Getting It Wrong&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;KYC/AML compliance&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Fragmented risk profiles, regulatory failures&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Fraud detection&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Missed patterns across linked identities&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;M&amp;amp;A data migration&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Duplicate customers in merged systems&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Regulatory reporting&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Incorrect unique customer counts&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;The cost of &lt;strong&gt;false negatives&lt;/strong&gt; (missed matches) is fragmented data and compliance risk. The cost of &lt;strong&gt;false positives&lt;/strong&gt; (incorrect matches) is merged records for different people—a data quality nightmare.&lt;/p&gt;
&lt;p&gt;Probabilistic matching lets you tune the tradeoff for your specific use case.&lt;/p&gt;
&lt;hr /&gt;
&lt;h2&gt;Key Takeaways&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Entity resolution&lt;/strong&gt; determines if records refer to the same real-world entity when no shared key exists&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Exact matching breaks&lt;/strong&gt; on real-world data—typos, format variations, and abbreviations defeat deterministic rules&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Probabilistic matching&lt;/strong&gt; weighs evidence across fields, producing a match score rather than a binary decision&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Common values&lt;/strong&gt; (María, García) provide weaker evidence than &lt;strong&gt;rare values&lt;/strong&gt; (Xiomara, Tlapoyawa)&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Blocking&lt;/strong&gt; reduces the comparison space from n² to something tractable&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Splink&lt;/strong&gt; handles the math, scales from laptops to clusters, and requires no labeled training data&lt;/li&gt;
&lt;/ul&gt;
&lt;hr /&gt;
&lt;h2&gt;Further Reading&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href="https://moj-analytical-services.github.io/splink/"&gt;Splink Documentation&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.robinlinacre.com/probabilistic_linkage/"&gt;Fellegi-Sunter Model Explained&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://link.springer.com/book/10.1007/978-3-642-31164-2"&gt;Data Matching (Springer)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://dataingovernment.blog.gov.uk/2022/09/23/splink-fast-accurate-and-scalable-record-linkage/"&gt;Record Linkage at Scale: A UK Government Perspective&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
</content:encoded>
    </item>
    <item>
      <title>Lineage IDs in Multimodal AI Pipelines</title>
      <link>https://alberto.codes/blog/2026-02-04-lineage-ids-in-multimodal-ai-pipelines</link>
      <guid isPermaLink="true">https://alberto.codes/blog/2026-02-04-lineage-ids-in-multimodal-ai-pipelines</guid>
      <pubDate>Wed, 04 Feb 2026 00:00:00 GMT</pubDate>
      <description>The simplest way to keep images, videos, model calls, and outputs tied together across retries and fan-out.</description>
      <content:encoded>&lt;p&gt;Multimodal pipelines look clean on a whiteboard:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;pull images + companion JSON&lt;/li&gt;
&lt;li&gt;render images → video&lt;/li&gt;
&lt;li&gt;send video + JSON to an AI API&lt;/li&gt;
&lt;li&gt;store outputs&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Then you ship it and reality shows up: retries, partial failures, parallel workers, version bumps, and the occasional &amp;quot;why does this output not match the input we think it does?&amp;quot;&lt;/p&gt;
&lt;p&gt;This is where lineage IDs earn their keep.&lt;/p&gt;
&lt;p&gt;If you’ve ever worked a dinner rush, it’s the same problem. Tickets keep printing, stations work in parallel, and sometimes you have to refire a dish. If your containers aren’t labeled and your tickets aren’t tracked, you end up staring at the pass thinking: &lt;strong&gt;is this plate actually for table 12?&lt;/strong&gt;&lt;/p&gt;
&lt;h2&gt;What “lineage” means (in practice)&lt;/h2&gt;
&lt;p&gt;Lineage is just &lt;strong&gt;the ability to answer “where did this come from?”&lt;/strong&gt; at every step.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Which image set produced this &lt;code&gt;video.mp4&lt;/code&gt;?&lt;/li&gt;
&lt;li&gt;Which &lt;code&gt;video.mp4&lt;/code&gt; + metadata produced this model response?&lt;/li&gt;
&lt;li&gt;Which response produced this &lt;code&gt;final.json&lt;/code&gt; we shipped downstream?&lt;/li&gt;
&lt;li&gt;If we retried, which attempt “won” and why?&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Queue dashboards and worker metrics tell you &lt;em&gt;what ran&lt;/em&gt;. Lineage tells you &lt;em&gt;what the run touched&lt;/em&gt; — the digital equivalent of a ticket number plus labels on every prep container.&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Rule of thumb: if you can’t answer “which exact bytes went into this model call?”, you don’t have lineage — you have hopeful logging.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h2&gt;The shape of a lineage-aware pipeline&lt;/h2&gt;
&lt;p&gt;This is what you're aiming for: a &lt;code&gt;run_id&lt;/code&gt; that propagates through workers, durable receipts that link inputs/outputs, and provider request IDs that let you correlate with the API vendor later.&lt;/p&gt;
&lt;p&gt;&lt;img src="https://alberto.codes/lineage-flow.svg" alt="Lineage flow through a multimodal pipeline" /&gt;&lt;/p&gt;
&lt;h2&gt;The minimum viable lineage model&lt;/h2&gt;
&lt;p&gt;Keep it boring. You only need a few identifiers.&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Name&lt;/th&gt;
&lt;th&gt;What it identifies&lt;/th&gt;
&lt;th&gt;Why it matters&lt;/th&gt;
&lt;th&gt;Kitchen analogy&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;run_id&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;One end-to-end workflow instance (created once at ingest)&lt;/td&gt;
&lt;td&gt;Correlates everything across fan-out and retries&lt;/td&gt;
&lt;td&gt;Ticket number&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;step&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;A named transformation stage (&lt;code&gt;fetch_images&lt;/code&gt;, &lt;code&gt;render_video&lt;/code&gt;, &lt;code&gt;call_ai&lt;/code&gt;, &lt;code&gt;persist_outputs&lt;/code&gt;)&lt;/td&gt;
&lt;td&gt;Lets you reason about progress and partial failures&lt;/td&gt;
&lt;td&gt;Station (prep, sauté, expo)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;attempt&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;The retry count for a step&lt;/td&gt;
&lt;td&gt;Distinguishes first-run vs refires&lt;/td&gt;
&lt;td&gt;Refire count&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;artifact_id&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;A stable identifier for bytes (often a content hash)&lt;/td&gt;
&lt;td&gt;Prevents “same path, different bytes” confusion&lt;/td&gt;
&lt;td&gt;Labeled batch / container&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;If you already use a task queue, you’ll also have a &lt;code&gt;task_id&lt;/code&gt; per task execution. That’s useful for ops, but it’s not a workflow lineage ID by itself.&lt;/p&gt;
&lt;h2&gt;Lineage vs idempotency vs tracing&lt;/h2&gt;
&lt;p&gt;These concepts are related, but they solve different problems:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Lineage&lt;/strong&gt; — &amp;quot;What produced this?&amp;quot; — artifacts and transformations across the whole workflow&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Idempotency&lt;/strong&gt; — &amp;quot;What if this runs twice?&amp;quot; — make retries and duplicates harmless&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Tracing&lt;/strong&gt; — &amp;quot;Where did the time go?&amp;quot; — timelines, spans, latency, bottlenecks&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;In practice, you'll often reuse the same &lt;code&gt;run_id&lt;/code&gt; as a correlation handle across all three, but keep the semantics separate:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Don't use a queue &lt;code&gt;task_id&lt;/code&gt; as your lineage key.&lt;/li&gt;
&lt;li&gt;Don't use an idempotency key as your run identifier (it's usually &lt;em&gt;step-scoped&lt;/em&gt;).&lt;/li&gt;
&lt;li&gt;Don't assume provider request IDs replace your own identifiers (they're necessary, not sufficient).&lt;/li&gt;
&lt;/ul&gt;
&lt;h2&gt;A file-first implementation (works before you have a DB)&lt;/h2&gt;
&lt;p&gt;You can get 80% of the value with a directory per run and a couple JSON “receipts”.&lt;/p&gt;
&lt;pre&gt;&lt;code class="language-text"&gt;runs/
  2026-02-04T18-41-12Z_d5c1.../          # run_id (timestamp + UUID, for easy sorting)
    manifest.json                        # one place to start
    inputs/
      frames/0001.jpg
      frames/0002.jpg
      metadata.json
    steps/
      010_fetch_images/
        receipt.json
      020_render_video/
        video.mp4
        receipt.json
      030_call_ai/
        request.json
        response.json
        receipt.json
      040_persist/
        outputs.json
        receipt.json
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;The only “rule” is: &lt;strong&gt;every step writes a receipt that declares its inputs and outputs&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;Here’s what a step receipt can look like:&lt;/p&gt;
&lt;pre&gt;&lt;code class="language-json"&gt;{
  &amp;quot;run_id&amp;quot;: &amp;quot;2026-02-04T18-41-12Z_d5c1f0b4e6b84aef9a7be8d07f2c3a1b&amp;quot;,
  &amp;quot;step&amp;quot;: &amp;quot;render_video&amp;quot;,
  &amp;quot;attempt&amp;quot;: 1,
  &amp;quot;started_at&amp;quot;: &amp;quot;2026-02-04T18:41:15Z&amp;quot;,
  &amp;quot;finished_at&amp;quot;: &amp;quot;2026-02-04T18:41:22Z&amp;quot;,
  &amp;quot;inputs&amp;quot;: [
    {&amp;quot;path&amp;quot;: &amp;quot;inputs/frames/0001.jpg&amp;quot;, &amp;quot;sha256&amp;quot;: &amp;quot;…&amp;quot;},
    {&amp;quot;path&amp;quot;: &amp;quot;inputs/frames/0002.jpg&amp;quot;, &amp;quot;sha256&amp;quot;: &amp;quot;…&amp;quot;},
    {&amp;quot;path&amp;quot;: &amp;quot;inputs/metadata.json&amp;quot;, &amp;quot;sha256&amp;quot;: &amp;quot;…&amp;quot;}
  ],
  &amp;quot;outputs&amp;quot;: [
    {&amp;quot;path&amp;quot;: &amp;quot;steps/020_render_video/video.mp4&amp;quot;, &amp;quot;sha256&amp;quot;: &amp;quot;…&amp;quot;, &amp;quot;media_type&amp;quot;: &amp;quot;video/mp4&amp;quot;}
  ]
}
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;That's lineage: a tiny, append-only audit trail you can reconstruct into a graph later.&lt;/p&gt;
&lt;p&gt;&lt;img src="https://alberto.codes/receipt-chain.svg" alt="Receipt chain linking artifacts via content hashes" /&gt;&lt;/p&gt;
&lt;p&gt;Each step's receipt links inputs to outputs via content hashes — forming a chain you can traverse in either direction.&lt;/p&gt;
&lt;h3&gt;Why content hashes beat &amp;quot;whatever the filename was&amp;quot;&lt;/h3&gt;
&lt;p&gt;When you’re debugging, the most painful failure mode is “file A got overwritten by a retry, and now A points to different bytes than it did yesterday.”&lt;/p&gt;
&lt;p&gt;In kitchen terms: someone reused the same unlabeled container, and now “sauce” could mean three different things depending on who touched it last.&lt;/p&gt;
&lt;p&gt;If your &lt;code&gt;artifact_id&lt;/code&gt; is a content hash (ex: SHA-256), you get:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Deduplication for free (same bytes → same ID)&lt;/li&gt;
&lt;li&gt;Cheap integrity checks (hash mismatch = corruption/overwrite)&lt;/li&gt;
&lt;li&gt;A stable join key across systems (filesystem, object storage, DB, logs)&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;You can still store by filename; just &lt;strong&gt;record the hash in the receipt&lt;/strong&gt; so you have a stable identifier.&lt;/p&gt;
&lt;p&gt;Here’s a tiny standard-library helper for hashing and run IDs:&lt;/p&gt;
&lt;pre&gt;&lt;code class="language-python"&gt;import hashlib
import uuid
from datetime import datetime, timezone
from pathlib import Path


def new_run_id() -&amp;gt; str:
    ts = datetime.now(timezone.utc).strftime(&amp;quot;%Y-%m-%dT%H-%M-%SZ&amp;quot;)
    return f&amp;quot;{ts}_{uuid.uuid4().hex}&amp;quot;


def sha256_file(path: str | Path) -&amp;gt; str:
    h = hashlib.sha256()
    with open(path, &amp;quot;rb&amp;quot;) as f:
        for chunk in iter(lambda: f.read(1024 * 1024), b&amp;quot;&amp;quot;):
            h.update(chunk)
    return h.hexdigest()
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;For large video files, hashing adds I/O overhead (you're reading every byte). For most pipelines the auditability is worth it — but if it becomes a bottleneck, faster non-cryptographic hashes like xxHash or BLAKE3 work fine for artifact IDs.&lt;/p&gt;
&lt;h2&gt;Propagating lineage through a task queue&lt;/h2&gt;
&lt;p&gt;The most important operational discipline is simple:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Generate &lt;code&gt;run_id&lt;/code&gt; once, then pass it to every task and log line.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;In task-queue terms, that usually means your task signature always includes &lt;code&gt;run_id&lt;/code&gt;, and every task writes its receipt under &lt;code&gt;runs/&amp;lt;run_id&amp;gt;/steps/&amp;lt;step&amp;gt;/…&lt;/code&gt;.&lt;/p&gt;
&lt;p&gt;Retries are expected in at-least-once systems (I wrote more about that in &lt;a href="https://alberto.codes/blog/2026-02-02-task-queues-idempotency-and-ai-pipelines"&gt;Task Queues, Idempotency, and AI Pipelines&lt;/a&gt;). Lineage doesn’t prevent duplicates — it makes duplicates understandable.&lt;/p&gt;
&lt;h2&gt;Structured logs (the breadcrumbs between receipts)&lt;/h2&gt;
&lt;p&gt;Receipts are your durable audit trail. Logs are your high-volume timeline.&lt;/p&gt;
&lt;p&gt;The simple rule: &lt;strong&gt;every log line emitted by workers should include &lt;code&gt;run_id&lt;/code&gt;, &lt;code&gt;step&lt;/code&gt;, and &lt;code&gt;attempt&lt;/code&gt;.&lt;/strong&gt; Then add fields that make debugging cheap, like:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;code&gt;task_id&lt;/code&gt; (from your queue)&lt;/li&gt;
&lt;li&gt;&lt;code&gt;artifact_id&lt;/code&gt; / &lt;code&gt;sha256&lt;/code&gt; for important inputs/outputs&lt;/li&gt;
&lt;li&gt;provider request/response IDs for model calls&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;If you use structured logging (for example &lt;code&gt;structlog&lt;/code&gt;), you’re aiming for events that look like this:&lt;/p&gt;
&lt;pre&gt;&lt;code class="language-python"&gt;import structlog


log = structlog.get_logger().bind(
    run_id=run_id,
    step=step,
    attempt=attempt,
    task_id=task_id,
)
log.info(&amp;quot;step_started&amp;quot;)
log.info(&amp;quot;step_finished&amp;quot;, output_sha256=output_sha256, duration_ms=duration_ms)
&lt;/code&gt;&lt;/pre&gt;
&lt;pre&gt;&lt;code class="language-json"&gt;{
  &amp;quot;event&amp;quot;: &amp;quot;step_finished&amp;quot;,
  &amp;quot;run_id&amp;quot;: &amp;quot;2026-02-04T18-41-12Z_d5c1f0b4e6b84aef9a7be8d07f2c3a1b&amp;quot;,
  &amp;quot;step&amp;quot;: &amp;quot;render_video&amp;quot;,
  &amp;quot;attempt&amp;quot;: 1,
  &amp;quot;output_path&amp;quot;: &amp;quot;steps/020_render_video/video.mp4&amp;quot;,
  &amp;quot;output_sha256&amp;quot;: &amp;quot;…&amp;quot;,
  &amp;quot;duration_ms&amp;quot;: 7421
}
&lt;/code&gt;&lt;/pre&gt;
&lt;h2&gt;Correlating AI API calls (OpenAI or otherwise)&lt;/h2&gt;
&lt;p&gt;When you call an AI API, you want to keep three IDs tied together:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;For every model call, tie together: &lt;strong&gt;your&lt;/strong&gt; &lt;code&gt;run_id&lt;/code&gt;, &lt;strong&gt;their&lt;/strong&gt; request ID, and the &lt;strong&gt;exact bytes&lt;/strong&gt; you sent and received.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;A concrete OpenAI example:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Send your own ID via the &lt;code&gt;X-Client-Request-Id&lt;/code&gt; request header (ASCII, ≤ 512 chars).&lt;/li&gt;
&lt;li&gt;Log/store the server-generated &lt;code&gt;x-request-id&lt;/code&gt; response header.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;A practical pattern is:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;code&gt;steps/030_call_ai/request.json&lt;/code&gt;: the request payload you sent (or a redacted version)&lt;/li&gt;
&lt;li&gt;&lt;code&gt;steps/030_call_ai/response.json&lt;/code&gt;: the raw response body you received&lt;/li&gt;
&lt;li&gt;&lt;code&gt;steps/030_call_ai/receipt.json&lt;/code&gt;: a summary that links inputs/outputs + provider IDs&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;When the API supports it, also include your identifiers in the request (for example: request metadata, a &lt;code&gt;user&lt;/code&gt; field, or custom headers in your own gateway). The goal is that support tickets and logs can be traced with either &lt;strong&gt;your&lt;/strong&gt; &lt;code&gt;run_id&lt;/code&gt; or &lt;strong&gt;their&lt;/strong&gt; request/response ID. OpenAI documents this under &lt;a href="https://platform.openai.com/docs/api-reference/authentication?api-mode=responses#debugging-requests"&gt;Debugging requests&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;For a provider call receipt, you’re usually trying to capture something like:&lt;/p&gt;
&lt;pre&gt;&lt;code class="language-json"&gt;{
  &amp;quot;run_id&amp;quot;: &amp;quot;2026-02-04T18-41-12Z_d5c1f0b4e6b84aef9a7be8d07f2c3a1b&amp;quot;,
  &amp;quot;step&amp;quot;: &amp;quot;call_ai&amp;quot;,
  &amp;quot;attempt&amp;quot;: 1,
  &amp;quot;provider&amp;quot;: &amp;quot;openai&amp;quot;,
  &amp;quot;client_request_id&amp;quot;: &amp;quot;run=2026-02-04T18-41-12Z_d5c1f0b4e6b84aef9a7be8d07f2c3a1b;step=call_ai;attempt=1&amp;quot;,
  &amp;quot;provider_request_id&amp;quot;: &amp;quot;…&amp;quot;,
  &amp;quot;provider_response_id&amp;quot;: &amp;quot;…&amp;quot;,
  &amp;quot;inputs&amp;quot;: [{&amp;quot;path&amp;quot;: &amp;quot;steps/020_render_video/video.mp4&amp;quot;, &amp;quot;sha256&amp;quot;: &amp;quot;…&amp;quot;}],
  &amp;quot;outputs&amp;quot;: [{&amp;quot;path&amp;quot;: &amp;quot;steps/030_call_ai/response.json&amp;quot;, &amp;quot;sha256&amp;quot;: &amp;quot;…&amp;quot;}]
}
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Even if you later move storage to S3/GCS and logs to Splunk/Datadog, that “receipt” pattern still holds. It’s just a join table you can grep.&lt;/p&gt;
&lt;h2&gt;What you get from doing this early&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Debugging that scales&lt;/strong&gt; — &amp;quot;this output is wrong&amp;quot; becomes &amp;quot;this run_id, this step, this exact artifact hash&amp;quot;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Safer retries&lt;/strong&gt; — you can see which attempt produced which bytes, and detect overwrites&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Replayability&lt;/strong&gt; — pick a &lt;code&gt;run_id&lt;/code&gt;, re-run starting at a step, and keep the old receipts for comparison&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Auditing&lt;/strong&gt; — you can answer &amp;quot;what did we send to the model?&amp;quot; without guessing&lt;/li&gt;
&lt;/ul&gt;
&lt;h2&gt;Common pitfalls&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;Treating a queue &lt;code&gt;task_id&lt;/code&gt; as a workflow ID (great for ops, terrible for lineage).&lt;/li&gt;
&lt;li&gt;Writing outputs to fixed paths without &lt;code&gt;attempt&lt;/code&gt; or hashes (retries overwrite history).&lt;/li&gt;
&lt;li&gt;Adding correlation IDs to &lt;em&gt;some&lt;/em&gt; steps but not all (you break the chain).&lt;/li&gt;
&lt;li&gt;Logging raw prompts/inputs instead of receipts + hashes (you create privacy risk and still can’t reliably join artifacts).&lt;/li&gt;
&lt;/ul&gt;
&lt;h2&gt;Key takeaways&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;Create a &lt;code&gt;run_id&lt;/code&gt; once at ingest and propagate it everywhere (tasks, logs, artifacts).&lt;/li&gt;
&lt;li&gt;Treat every step like a station: write a receipt that names exact inputs and outputs.&lt;/li&gt;
&lt;li&gt;Prefer content hashes for &lt;code&gt;artifact_id&lt;/code&gt; so retries and overwrites don’t blur history.&lt;/li&gt;
&lt;li&gt;Capture provider request/response IDs for model calls so you can correlate later.&lt;/li&gt;
&lt;li&gt;Start file-first; you can always move the same receipts into a DB or a catalog later.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2&gt;Closing thought&lt;/h2&gt;
&lt;p&gt;You don't need a data catalog or a dedicated lineage platform to start. A &lt;code&gt;run_id&lt;/code&gt;, a couple receipts, and disciplined propagation through your tasks will save you days of debugging — and a lot of unnecessary model spend — once your pipeline hits the dinner rush.&lt;/p&gt;
</content:encoded>
    </item>
    <item>
      <title>Task Queues, Idempotency, and AI Pipelines</title>
      <link>https://alberto.codes/blog/2026-02-02-task-queues-idempotency-and-ai-pipelines</link>
      <guid isPermaLink="true">https://alberto.codes/blog/2026-02-02-task-queues-idempotency-and-ai-pipelines</guid>
      <pubDate>Mon, 02 Feb 2026 00:00:00 GMT</pubDate>
      <description>Why at-least-once delivery means your AI pipeline will process duplicates, and why idempotency is the only reliable fix.</description>
      <content:encoded>&lt;p&gt;If you run AI workloads through a task queue, duplicates are not a bug. They are a feature of the delivery guarantee you almost certainly chose. Understanding why they happen, and designing for them, is the difference between a pipeline that wastes money on redundant LLM calls and one that shrugs off failures gracefully.&lt;/p&gt;
&lt;p&gt;Think of it like a busy kitchen. A ticket printer spits out orders, and line cooks pull them. If a cook drops a plate halfway through and nobody crosses the ticket off, the expeditor calls it again. The kitchen makes the dish twice. That's at-least-once delivery: better to send an extra plate to the pass than to lose an order entirely.&lt;/p&gt;
&lt;h2&gt;Why Exactly-Once Delivery Is So Rare&lt;/h2&gt;
&lt;p&gt;Most engineers want exactly-once delivery: every message processed one time, no more, no less. The problem is that exactly-once requires coordination between the broker and the worker that is extremely hard to achieve in practice. The broker needs to know the worker finished, and the worker needs to know the broker won't send the message again. In a distributed system where either side can crash, restart, or lose a network connection at any moment, that coordination breaks down.&lt;/p&gt;
&lt;p&gt;What you actually get in most systems is one of two alternatives:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;At-most-once&lt;/strong&gt;: the broker deletes the message before the worker confirms success. If the worker crashes, the message is gone. Simple, but you lose work.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;At-least-once&lt;/strong&gt;: the broker keeps the message until the worker explicitly acknowledges it. If anything goes wrong, the message gets redelivered. You never lose work, but you get duplicates.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;At-least-once is the default in SQS, Celery with &lt;code&gt;acks_late=True&lt;/code&gt;, Google Cloud Tasks, and most durable workflow engines like Temporal. It's the pragmatic choice for pipelines where losing a task is worse than doing it twice.&lt;/p&gt;
&lt;h2&gt;How Duplicates Actually Happen&lt;/h2&gt;
&lt;p&gt;Duplicates are not some theoretical edge case. Three scenarios produce them regularly:&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Worker crash mid-processing.&lt;/strong&gt; A worker picks up a message, starts an expensive OCR extraction, and gets OOM-killed halfway through. The broker never received an acknowledgment, so it redelivers the message to another worker. Now two attempts exist: one partial (dead), one fresh.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Visibility timeout expiry.&lt;/strong&gt; In SQS, when a worker receives a message, that message becomes invisible to other consumers for a configurable window (the visibility timeout). If your AI pipeline task takes longer than expected—say a large PDF triggers a slow embedding model—the timeout expires and SQS hands the message to a second worker while the first is still running.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Late acknowledgment.&lt;/strong&gt; The worker finishes the job and sends an ack, but the network blips. The broker assumes the worker failed and redelivers. The work was already done, but the queue doesn't know that.&lt;/p&gt;
&lt;p&gt;All three scenarios result in the same task being executed more than once.&lt;/p&gt;
&lt;h2&gt;Idempotency: Making Duplicates Harmless&lt;/h2&gt;
&lt;p&gt;Idempotency means that running an operation multiple times produces the same result as running it once. Back to the kitchen: if you've already plated the risotto and it's sitting under the heat lamp, you don't fire a second one just because the ticket printed again. You check the pass first. Same idea in code.&lt;/p&gt;
&lt;p&gt;The most common pattern is an &lt;strong&gt;idempotency key&lt;/strong&gt;—a unique identifier derived from the input that you check before doing expensive work:&lt;/p&gt;
&lt;pre&gt;&lt;code class="language-python"&gt;def process_document(task: dict):
    idempotency_key = f&amp;quot;doc:{task['document_id']}:v{task['version']}&amp;quot;

    if store.exists(idempotency_key):
        return  # Already processed, skip

    chunks = split_and_embed(task[&amp;quot;content&amp;quot;])  # Expensive LLM call
    vector_db.upsert(task[&amp;quot;document_id&amp;quot;], chunks)

    store.set(idempotency_key, ttl=86400)  # Mark as done
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;The key is derived from the document ID and version so that a legitimate re-processing (new version) still runs, but a duplicate delivery of the same version gets caught. The TTL on the key keeps your idempotency store from growing forever.&lt;/p&gt;
&lt;p&gt;This is not clever. It's a few lines of code. But it's the difference between paying for one embedding call and paying for three because a worker got slow under load.&lt;/p&gt;
&lt;h2&gt;The Redelivery Lifecycle&lt;/h2&gt;
&lt;p&gt;When you combine at-least-once delivery with retry policies and a dead-letter queue, message flow looks like this:&lt;/p&gt;
&lt;p&gt;&lt;img src="https://alberto.codes/redelivery-lifecycle.svg" alt="Redelivery lifecycle: task flows from queue to worker, retries on failure, and lands in a dead-letter queue after the retry limit" /&gt;&lt;/p&gt;
&lt;p&gt;The dead-letter queue (DLQ) is not where messages go to die. It's a &lt;strong&gt;throughput protection mechanism&lt;/strong&gt;. Without it, a poison message—a malformed PDF, an input that consistently crashes your OCR service—cycles through the queue forever, consuming worker capacity and blocking healthy tasks behind it. The DLQ catches these after a configured number of retries and moves them aside so the rest of the pipeline keeps flowing. You review them later, fix the root cause, and replay if needed.&lt;/p&gt;
&lt;p&gt;Every kitchen has an 86 board—the list of items that are cut from service because you're out of an ingredient or a piece of equipment is down. You don't keep trying to plate a dish you can't finish. You pull it and deal with it. A DLQ is your 86 board for tasks that can't be completed right now.&lt;/p&gt;
&lt;h2&gt;Why Retries Cost More in AI Pipelines&lt;/h2&gt;
&lt;p&gt;In a traditional web backend, retrying a database insert is cheap. In an AI pipeline, retries hit your wallet directly.&lt;/p&gt;
&lt;p&gt;Consider a document ingestion pipeline: receive a PDF, OCR it, chunk it, generate embeddings, then extract structured fields with an LLM. If extraction fails and the task retries from the top, you're paying for OCR and embeddings again. At scale, this adds up:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Embedding generation&lt;/strong&gt;: a 50-page document might produce 200 chunks at ~500 tokens each. One unnecessary retry means 100,000 extra tokens through your embedding model.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;LLM extraction&lt;/strong&gt;: a structured extraction call with a long context window can cost $0.05–$0.20 per call. Ten retries across a batch of 1,000 documents is an extra $500–$2,000.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Tool calls&lt;/strong&gt;: if your agent pipeline triggers external APIs (geocoding, enrichment, web search) on every retry, you're multiplying both cost and rate-limit pressure.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Idempotency keys at each stage let you skip work that already completed successfully. The OCR result is cached. The embeddings are already in the vector store. The retry only re-executes the extraction step that actually failed. This is sometimes called &lt;strong&gt;partial progress recovery&lt;/strong&gt;, and it turns an expensive full retry into a cheap targeted one.&lt;/p&gt;
&lt;h2&gt;Key Takeaways&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;At-least-once delivery is the default in most task queue systems. Duplicates are expected, not exceptional.&lt;/li&gt;
&lt;li&gt;Idempotency keys are cheap to implement and directly reduce wasted LLM and embedding spend on retries.&lt;/li&gt;
&lt;li&gt;Dead-letter queues protect pipeline throughput by isolating poison messages instead of letting them loop forever.&lt;/li&gt;
&lt;li&gt;AI pipelines should cache intermediate results (OCR output, embeddings) so retries only redo the step that failed.&lt;/li&gt;
&lt;li&gt;Design your idempotency keys from the input data (document ID + version), not from internal state that might change between attempts.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2&gt;Further Reading&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href="https://docs.aws.amazon.com/AWSSimpleQueueService/latest/SQSDeveloperGuide/sqs-visibility-timeout.html"&gt;Amazon SQS Visibility Timeout&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://docs.aws.amazon.com/AWSSimpleQueueService/latest/SQSDeveloperGuide/sqs-dead-letter-queues.html"&gt;Amazon SQS Dead-Letter Queues&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://docs.celeryq.dev/en/stable/userguide/tasks.html"&gt;Celery Task Acknowledgment and acks_late&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://docs.temporal.io/evaluate/understanding-temporal"&gt;Temporal: Understanding Durable Execution&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://docs.hatchet.run/home"&gt;Hatchet – Distributed Task Queue for Python&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.oreilly.com/library/view/designing-data-intensive-applications/9781491903063/ch11.html"&gt;Designing Data-Intensive Applications: Chapter 11 – Stream Processing (O'Reilly)&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
</content:encoded>
    </item>
    <item>
      <title>Why I Chose Reflex for My Portfolio Site</title>
      <link>https://alberto.codes/blog/2026-02-01-why-reflex-for-a-python-engineers-portfolio</link>
      <guid isPermaLink="true">https://alberto.codes/blog/2026-02-01-why-reflex-for-a-python-engineers-portfolio</guid>
      <pubDate>Sun, 01 Feb 2026 00:00:00 GMT</pubDate>
      <description>A Python engineer's case for building a portfolio site without touching JavaScript.</description>
      <content:encoded>&lt;p&gt;When I decided to build a personal site, the default answer was obvious: grab a React template, maybe Next.js, throw it on Vercel, and call it a day. That's what most developer portfolios run on. But I'm not a frontend engineer. I'm a Python engineer who works on AI systems, and I wanted my site to reflect that.&lt;/p&gt;
&lt;p&gt;So I built it with &lt;a href="https://reflex.dev"&gt;Reflex&lt;/a&gt;.&lt;/p&gt;
&lt;h2&gt;What Is Reflex?&lt;/h2&gt;
&lt;p&gt;Reflex is a Python framework for building full-stack web applications. You write Python — components, state management, routing, everything — and Reflex compiles it to a React frontend with a FastAPI backend. No JavaScript required (unless you want it).&lt;/p&gt;
&lt;p&gt;Here's what a simple page looks like:&lt;/p&gt;
&lt;pre&gt;&lt;code class="language-python"&gt;import reflex as rx

def hello() -&amp;gt; rx.Component:
    return rx.container(
        rx.heading(&amp;quot;Hello, world&amp;quot;),
        rx.text(&amp;quot;Built with Python.&amp;quot;),
    )
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;That's real, working code. No JSX, no &lt;code&gt;useState&lt;/code&gt;, no &lt;code&gt;npm install&lt;/code&gt;. Just Python.&lt;/p&gt;
&lt;h2&gt;Why Not Next.js?&lt;/h2&gt;
&lt;p&gt;There's nothing wrong with Next.js. It's battle-tested, well-documented, and has a massive ecosystem. But for me, it came down to a few things:&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Context switching is expensive.&lt;/strong&gt; My day job is Python — agents, pipelines, data engineering. Dropping into TypeScript and the React ecosystem for a side project meant learning (or re-learning) a different set of patterns, tooling, and conventions. With Reflex, I stay in the language I think in.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;I don't need the React ecosystem.&lt;/strong&gt; A portfolio site doesn't need server components, incremental static regeneration, or a component library with 200 options. It needs a few pages, some text, and maybe a blog. Reflex handles that without the overhead.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;The site itself is a signal.&lt;/strong&gt; If I'm positioning myself as a Python and AI engineer, it makes sense for the site to be built with Python. It's a small thing, but it's consistent.&lt;/p&gt;
&lt;h2&gt;What Reflex Gets Right&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Components feel natural.&lt;/strong&gt; If you've written Python, you can read Reflex code. Components are functions that return component trees. Styling uses keyword arguments. There's no template language to learn.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;State management is simple.&lt;/strong&gt; Reflex uses Python classes for state. You define variables, write event handlers as methods, and bind them to components. It's closer to how you'd think about state in a backend application.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Deployment is straightforward.&lt;/strong&gt; &lt;code&gt;reflex run --env prod&lt;/code&gt; gives you a production build. You get a FastAPI server you can deploy anywhere Python runs — no Node.js runtime needed in production.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;It includes what you need.&lt;/strong&gt; Routing, SEO meta tags, responsive breakpoints, markdown rendering, sitemaps — these come built in or as first-party plugins. I didn't have to research and install ten packages to get a basic site working.&lt;/p&gt;
&lt;h2&gt;What to Watch Out For&lt;/h2&gt;
&lt;p&gt;Reflex isn't perfect, and I'd be doing you a disservice if I didn't mention the tradeoffs:&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Smaller community.&lt;/strong&gt; When you hit an issue, there are fewer Stack Overflow answers and blog posts to lean on compared to React or Next.js. The Discord community is active, but you'll sometimes need to read source code.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Component library is growing but limited.&lt;/strong&gt; You won't find the equivalent of every Radix or shadcn component. Reflex wraps Radix Themes and gives you a solid set of primitives, but if you need something niche, you may need to write a custom component or wrap a React library yourself.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Performance ceiling.&lt;/strong&gt; For a portfolio site, performance is fine. For a complex, highly interactive application with thousands of concurrent users, you'd want to evaluate more carefully. The Python backend handles state, which adds a round-trip that a pure client-side React app avoids.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;It's still maturing.&lt;/strong&gt; Reflex is under active development. APIs occasionally change between versions. This is improving, but it's worth noting if stability is your top priority.&lt;/p&gt;
&lt;h2&gt;Who Should Consider Reflex?&lt;/h2&gt;
&lt;p&gt;If you're a Python developer building something like:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;A personal site or portfolio&lt;/li&gt;
&lt;li&gt;An internal tool or dashboard&lt;/li&gt;
&lt;li&gt;A prototype or MVP&lt;/li&gt;
&lt;li&gt;A data-driven app where Python is already doing the heavy lifting&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Reflex is worth a serious look. You'll ship faster by staying in one language, and the result is a real web app — not a Streamlit dashboard with limitations.&lt;/p&gt;
&lt;h2&gt;The Bottom Line&lt;/h2&gt;
&lt;p&gt;I chose Reflex because it let me build a professional site using the tools I already know, without compromising on the result. The site is fast, responsive, and does everything I need. And when I want to add features — like the blog you're reading right now — I'm writing Python, not context-switching into a different ecosystem.&lt;/p&gt;
&lt;p&gt;If you're a Python engineer who's been putting off building a personal site because you don't want to deal with JavaScript tooling, give Reflex a try.&lt;/p&gt;
</content:encoded>
    </item>
  </channel>
</rss>
