CounterProof

The Number We Couldn’t Back Up

An early draft of one of our own advisories cited a reproduction figure with no surviving artifact behind it. Here is what caught it, what we did about it, and why a withdrawal is the strongest evidence a review method can produce.

Most of what a review practice says about its own rigour is a claim. This is a receipt.

In August 2026, while preparing an internal advisory on a consensus-liveness issue in an open-source Bitcoin protocol we already work in, an early draft stated a reproduction figure for one of the findings — a specific fraction, cited as measured. It read as evidence. It was not: no harness, log file, or branch existed anywhere that could reproduce, or even locate, where that number had come from. It had been asserted, not measured, and it had survived into a draft advisory under the name of a practice whose stated rule is no unmeasured magnitude.

Our own instrument-defect catalogue — the internal list we keep of ways a verification process can quietly report a false result — caught it. Before the advisory went anywhere near a client or a maintainer, an internal check asked the plain question every claim in an assessment has to survive: where is the artifact behind this number? There wasn’t one.

What we did about it

What happened next is the actual point of this note, more than the mistake itself. We rebuilt the test from scratch, measured what could honestly be measured — a persistent-attacker scenario reproduced 20 times out of 20 attempts, with a related but distinct scenario left explicitly not measured rather than assumed — and rewrote the advisory to withdraw the earlier figure in writing, on the record, before the corrected version went anywhere. Not softened. Not quietly replaced. Withdrawn, with the reason stated alongside it.

That is a small, unglamorous event, and we are not telling you about it because it is a triumph. We are telling you about it because it is the only kind of evidence that actually means anything for a practice whose entire pitch is that assessments should be evidence-graded and inspectable rather than trusted on reputation. Anyone can write on a website that their process catches errors. Very few will show you the specific error it caught in their own work, on the record, with what changed as a result. A method that has never found a defect in its own output has not been tested — it has just not yet been unlucky.

The rule that came out of it

The standing rule is simple, and it is now permanent in how we work: a number in an assessment is either backed by an artifact a second party could independently check, or it is not in the assessment. Confidence is not a substitute for a file on disk. We would rather publish fewer numbers, later, than one we cannot hand you the receipt for.

It is also why every CounterProof engagement carries a standing withdrawal contract: if a finding we deliver you is later shown wrong, we correct it in writing, the same way, no differently than we did here on our own work before it ever had the chance to matter for a client.


CounterProof is an independent adversarial review practice of CounterProof Research Limited (Malta). Proof attests; CounterProof refutes. counterproof.io

← All notes