Notes

The instrument lies before the model does

Shane Pilon · 2026-07-05

For months, a capability I was building measured zero. Not low — zero. I built increasingly elaborate theories about why, and increasingly elaborate mechanisms to fix it.

The evaluation harness loaded checkpoints with strict=False. It was silently dropping 112 trained keys. Every one of those evaluations tested a model whose routing was random.

The capability had never been the problem. The instrument had been broken the entire time, and it failed in the most expensive possible direction: it produced a number, consistently, that was wrong.

What I took from it is the first of the four habits the work runs on, and it is the first step of any evaluation I run or review. The longer version, with the attack list and the fix, is in the sample teardown.

This is the format an engagement delivers in. See what you can buy, or send me one number you do not trust.

Subscribe by RSSAll notes