Notes
Written by Shane Pilon. Findings from my own evaluation harness, in the format a client engagement delivers them in. Three notes so far, and two of them are one incident at two lengths: the short note, and the sample teardown it points at.
2026-08-06
Never branch on a tensor value you cannot reproduce
A Python branch that reads a value off a tensor is a branch floating-point noise can decide. One of those killed a training run of mine, loudly, because a checkpoint sanity check happened to be counting. That crash is the only instance I have. Here is the grep that finds them.
2026-08-02
Read the code that made the number
A sample teardown, run on my own harness so you can see the format before you buy one. One capability read zero for months. The metric as received, the attack list in order, the finding, the steps to reproduce it, and what the number was actually worth once the confound was gone.
2026-07-20
The honesty floor was free
Under the evaluation harness, on a thin 21-question multi-turn benchmark of my own with no external baseline, making memory fabrication structurally impossible cost zero performance. Identical scores, seed for seed.
2026-07-05
The instrument lies before the model does
Months of zero-percent results were a loader flag, not a capability ceiling.
If a number in your own harness is about to carry weight it has not carried before, send me the number, or read what you can buy.
