The instrument lies before the model does
Shane Pilon · 2026-07-05
For months, a capability I was building measured zero. Not low — zero. I built increasingly elaborate theories about why, and increasingly elaborate mechanisms to fix it.
The evaluation harness loaded checkpoints with strict=False. It was silently
dropping 112 trained keys. Every one of those evaluations tested a model whose
routing was random.
The capability had never been the problem. The instrument had been broken the entire time, and it failed in the most expensive possible direction: it produced a number, consistently, that was wrong.
What I took from it is the first of the four habits the work runs on, and it is the first step of any evaluation I run or review. The longer version, with the attack list and the fix, is in the sample teardown.
This is the format an engagement delivers in. See what you can buy, or send me one number you do not trust.