CounterProof

We withdrew our own number

We shipped a reproduction figure that had no retained artifact behind it. We found it, withdrew it in writing, and replaced it with a narrower claim that is actually measured. This is the withdrawal contract being exercised — on us.

Every report we issue carries a standing commitment: if a finding is shown wrong, it is retracted in writing. That sentence is easy to publish and cheap to mean until the first time it costs you something.

On 11 August 2026 it cost us something, and the wrong claim was ours.

What we shipped

An advisory of ours reported a reproduction figure for a finding: the attack succeeded 18 of 20 times — 90%. It sat in the report as a reliability measurement, and it carried an evidence grade that assumed a measurement had been made.

What was wrong with it

When we went back to the artifact behind that number, there wasn’t one. No harness, no logs, no branch. Nothing on the machine reproduced it. The figure was real in the sense that someone had run something and remembered a result; it was not real in the sense that anyone — including us — could go and check.

It was also a single-window number describing a specific attack timing, presented without that qualifier, which made it read as a broader claim than the underlying work supported even if the artifact had existed.

Our own rules, published before this happened, say no unmeasured magnitude and every claim survives cross-lineage attack before it ships. An evidence-graded number resting on nothing backed violates both. It got through anyway.

What we did

We withdrew it in writing, in the advisory itself, at the top, rather than quietly editing the number and reissuing. The withdrawal is part of the document’s permanent record. A reader who comes to it later sees that a figure was claimed, that it failed, and what replaced it.

Then we rebuilt the measurement properly: a fresh harness, retained this time, run in a properly instrumented environment. The persistent case succeeded against the target 20 out of 20 times, with 5 of 5 clean controls — controls being the runs that must not trigger the failure, without which a 20/20 result means nothing.

What we did not claim

This is the part that matters more than the new number.

The rebuilt harness measures the persistent attacker. It does not measure the single-window, per-attempt success rate — which is what the withdrawn 18/20 had purported to describe. So we state, explicitly, that the per-attempt rate was not measured and no figure is claimed for it.

That leaves the report with a narrower claim than the one it started with, and a stronger one where it still stands. A reader gets a measured number for the thing we measured and an honest blank where we didn’t. The temptation — obvious, and the reason most reports don’t read like this — is to let the new 20/20 quietly occupy the space the old 18/20 vacated, since both are numbers about the same finding and nobody would have noticed. They measure different things. Saying so costs a sentence and buys the only thing an assessment is actually for.

Why we are publishing this

Two reasons, and neither is modesty.

The first is that a withdrawal contract nobody has ever seen exercised is a marketing line. Once exercised, in public, on the practice’s own evidence, it is a demonstrated property. You now know what happens when we are wrong, because it has happened and the record is there. That is not a claim about our carefulness; it is a fact about our conduct that you can go and check.

The second is that this is the exact failure our own work is about. We argue that machine-assisted review has a characteristic defect — fluent, well-formed output that is not backed by anything, arriving in the same voice as output that is. An evidence-graded figure with no artifact behind it is precisely that defect, committed by the people who wrote the catalogue of it. Nobody is outside this problem. The question is only whether your process catches it, and whether you say so when it does.

We would rather be the practice that publishes its own retraction than the one that has never needed to.


The advisory containing this withdrawal is a client- and maintainer-facing document and is not reproduced here; the withdrawal, the replacement measurement, and the explicit not-measured statement are all recorded in it. If we get something wrong on this page, we will correct it here, in writing, and say what changed.

← All notes