Notes

The honesty floor was free

Shane Pilon · 2026-07-20

The standard assumption is that safety properties cost capability. You constrain the model, and it gets worse at its job. The constraint is a tax you pay for trustworthiness.

I measured it. The tax was zero.

The configuration with a hard mask over the recall path scored 57.1. The unmasked configuration scored 57.1. Identical, seed for seed, across two seeds.

Two things travel with those numbers wherever they go. Both come from my own multi-turn memory benchmark, which is an internal measure with no external baseline, and the result is scoped to that multi-turn axis and nowhere else. And the instrument is a 21-question set, which is thin, and I have said so in my own notes.

The scope has to travel with that too, so here it is. The mask only reads a flag that the evaluation harness sets, and the harness can set it because it knows what was stored. The deployed chat path never arms it. So the property this measures is "unstored content is unavailable on the recall path under the harness", not "the shipped model cannot fabricate on free-form input". The second one is unsolved.

That reframes the property entirely. It is not a tradeoff to be justified. It is a free structural choice, which means the only reason not to build it in from the start is not having thought of it.

The caveat I would apply to my own result: this is one architecture, one metric, two seeds. It says the tax can be zero. It does not say it always is.

This is the format an engagement delivers in. See what you can buy, or send me one number you do not trust.

Subscribe by RSSAll notes