Alberto Nieto

Hello, I'm

Alberto Nieto

Generative AI Principal Engineer

I started as a teller, taught myself to code, and spent 26 years building my way to Principal Engineer. Now I design AI systems at work, and on my own time I build open source tools that test what models actually do, then publish the evidence.

2x Top Performer

2 Patents Pending

26 Years

8 PyPI Packages

Latest writing

All posts →

Gemma 4 read a love note as a scam. A rewritten question brought it within two messages of Jev.

I asked Jev and Gemma 4 31B the same scam question about 158 public text messages. Jev got 151 right. Gemma, on my own card, got 131, because it flagged love notes and everyday chat as fraud. Then I let an optimizer rewrite the question once for each model. Gemma climbed to 153, and on calibration its percentages ended closer to reality than Jev's.

My Gemma was 99.96 percent sure the total matched. It had not seen the receipt.

typevet is a small Python library, public today, that asks an open model typed questions and reads the answer as a probability. I asked Gemma 4 31B, on my own 24 GiB card, whether an expense claim matched a real receipt. Without the photo it said yes to every claim that showed a full number, almost certain every time. With the photo it caught all six wrong totals.

Jev looked underconfident on Banking77. Two of my labels were the reason.

I asked TypeSafe's Jev one yes-or-no question about a thousand public messages and checked whether its percentages matched what actually happened. On scam text messages they did. On bank support messages they looked far too cautious, until I read the messages themselves. The problem was two of my labels, not Jev.

Areas of Expertise

  • Generative AI & Agents
  • Python
  • Data Engineering & Pipelines
  • Cloud (GCP, PCF)
  • OCR & Document Intelligence
  • CI/CD & DevOps
  • Enterprise Architecture
  • Technical Leadership

© 2026 Alberto Nieto. All rights reserved.