In August 2026 we ran the experiment our own case report had been owing for months: does the full review apparatus we built earn its cost against one cheap pass? We pre-registered it, scored it blind, and used published professional audits as the answer key. It lost one target outright and could not be named the winner on the other. Then, in September, a retrospective audit of the runs found that the comparison was not as clean as we had said. This note carries all of it, because the rule this practice sells (a practice that cannot show you where it was wrong has no standing where it says it is right) applies to us first.
What we did
Two targets, both public code with a public professional audit at a stated commit, so that the answer key existed before we started and was not ours:
- Target 1, BDK (
bitcoindevkit/bdk_walletat775e4aee, v1.0.0-beta.2), against the Wizardsardine audit of December 2024: twelve findings as the denominator. Run 14 August. - Target 2, ZF FROST (
ZcashFoundation/frostat5fa17ed, v0.6.0,frost-core), against the NCC Group assessment of October 2023: four findings as the denominator. Run 22–23 August.
Two arms on each. Arm A, the cheap tier: one pass by a single outside model lineage with an inline citation self-check. Arm B, the full apparatus: a panel of finder lineages, then a “designer’s advocate” seat that argues the code’s intended trust assumptions, then an adversarial reachability check, then a gate. Both arms worked from the pre-fix source with the same finding-agnostic rules of engagement and were never told what the audit had found. Denominators and the match rule (same sink and same mechanism) were frozen in writing before scoring. Each report was then scored by a blind analyst that had run neither arm and verified every claim against the source. The keep-rule was pre-committed: the apparatus earns its cost if it finds at least two more of the auditor’s findings, or if the cheap arm ships at least two more source-refuted claims.
What we found
| Cheap tier | Full apparatus | |
|---|---|---|
| Target 1, recall of the auditor’s 12 findings | 8 | 6 (5 under a strict reading) |
| Target 1, source-refuted claims shipped | 0 of 25 | 1 of 39 |
| Target 2, recall of the auditor’s 4 findings | 4 | 0–1 (scorer-dependent, see below) |
| Target 2, source-refuted claims shipped | 3 of 11 | 1 of 7 (0 under the second scorer) |
On the first target the apparatus was simply worse: lower recall, one more false claim, roughly six times the model calls, two seat outages. On the second it did the one thing our hypothesis had predicted, cutting false claims in a noisy setting (its pre-registered precision trigger fired) and bought it by deleting the findings the auditor had ruled real (all four under one scorer, three of four under another), after its own finder panel had found all four. The pre-registered rule had no way to weigh a precision gain against a recall collapse, so it could not name a winner; we do not call that a win.
The mechanism was the same on both targets and it was not the finders. The advocate and reachability seats remove any candidate that rests on “an assumption the code documents, or a party the threat model trusts”. That test is satisfied both by false positives (safe to remove) and by real findings whose whole point is that the documented assumption is not enforced. On FROST the code’s own comments said the caller validates the inputs; the auditor had ruled that exactly this allocation was the defect; our advocate believed the comments. The individual dismissals were defensible readings. Settling the question silently, by deletion, was the error.
What we changed
- Filters annotate; they no longer delete. A seat that would drop a finding on assumption, intent or trust grounds now flags it and leaves it in the report for a person to adjudicate.
- We re-assembled target 2’s retained verdicts in that mode, against a pre-registered test (23 August; the filters’ existing verdicts became flags, a labelling change rather than a fresh run, scored by one blind scorer over all 24 candidates this time). Recall returned to 4 of 4, mechanically, since nothing was deleted. The interesting number was the flag: every one of the seven false claims carried a flag, and none of the seven unflagged items was false, on this one target the flag worked as a screen. But ten of the seventeen flagged items were real, six of them the auditor’s own. As a verdict the same flag is destructive. The flag tells you what to look at, not what to throw away: and what it costs a reviewer to adjudicate seventeen flagged items, or whether a person under deadline resists deleting on a flag, is untested.
- The cheap tier is the default: a policy choice on two runs, provisional, and kept after the audit below narrowed the comparison to its precision delta. The heavy panel is reserved for a surface with a genuinely high raw false-positive rate where a person will adjudicate the flags.
- A “precision win” no longer counts on its own. On target 2 the precision threshold fired because the same filter had cratered recall; the pre-registration had treated the two as independent dials, and the first draft of our own write-up went on to invent a weighting after seeing the result, a review gate caught it. Forward, a precision gain counts only if the apparatus’s recall is no more than one finding below the cheap arm’s.
Then the measurement failed its own audit
On 13 September a retrospective audit of every stored seat session (1,385 of them) found that the cheap arm was not as blind as the protocol had said. Its sessions had fetched, among other pages, the targets’ own issue trackers: issues naming the very functions and objects behind several of the auditor’s findings. None of those pages states the audit’s mechanisms, so nobody read an answer key; but the protocol’s sentence that the arms “see only the pinned source” was false, and the contamination control we had run was a cold memory probe, which cannot detect retrieval at run time. The apparatus arm’s sessions had been purged by its own driver, so its exposure is unknowable.
A same-day follow-up went finding by finding. On BDK, two of the eight rows credited to the cheap arm sit on the function named in the fetched issues; the other six show no captured overlap. On FROST, three of the five credited rows sit on the exact parameter the one fetched issue is about: a direct topical hit, not merely reachability, though the issue never states the mechanism. Even crediting the cheap arm with only the uncorrelated rows (six of eight, two of five) it still recalled at least as much as the apparatus on both targets, so the direction stands; the margin does not. And the follow-up could not be finished: the cheap arm’s session logs had been pruned by the tool’s own housekeeping within hours of the first audit, and its raw finding text had never been saved, so the full fetch lists and the per-finding timing are unrecoverable.
What that overturns: the cheap tier’s recall figures as clean numbers. What it does not overturn: that the measurement happened, that the oracles were external, that pre-registration and blind scoring were real, and (the finding that matters) that the apparatus’s filter seats deleted findings a professional auditor had ruled real. That was established by reading the seats’ own dismissals, not by the recall counts. Both audits were run by the same model vendor as the orchestrator; they are on the record with that label.
Five more weaknesses the record carries, most of them found by review seats rather than by us: the blind scorer was from the same model vendor as the orchestrator and as the apparatus’s merge seat, forced by availability and mitigated by source verification, but a correlation channel in an experiment about decorrelation; the same outside lineage that ran the cheap arm also sat as the apparatus’s reachability verifier on both targets, so part of the recall gap is one lineage disagreeing with itself across roles; the cheap arm as run was thinner than the one pre-registered (no separate tracker check or control suite), and on target 1 the memory probe’s result was never written to the run log; the apparatus ran with two finder lineages against a pre-registered minimum of three, so a full-strength version might do better, though the losses came from the filters, which a third finder would not change; and a third blind scorer, re-scoring the seven apparatus findings on target 2, agreed on six and flipped the load-bearing one, which is why that cell reads 0–1 and not 0.
What this buys you
Nothing on this page is a catch-rate claim; it is the opposite. What it is: the capability to run this measurement, pre-registered, blind-scored, against a public audit, with a contamination check that inspects what a seat fetched rather than what it remembers, and copies the session log to a durable store at scoring time because the logs behind this audit were gone within the day, is adopted in the same revision of our case report that publishes this measurement, and the first thing it measured was us. We did not find, when we searched the market in September 2026, another review practice that publishes a controlled comparison of its own product, let alone one it lost. That is a statement about what is on the record, not about who is better; the next cheap pass may beat us again, and if it does you will read it here.
Limits that belong here: two targets, one of them with a four-row denominator; the cheap arm’s recall figures are retrieval-contaminated and the apparatus arm’s exposure is unknowable; the scorer was same-vendor; costs were counted in model calls, which flatters the cheap tier (its single pass on target 1 was the most expensive measured call); the per-finding analysis of which cheap-arm hits followed a fetched page is partial and can no longer be completed. Our case report’s statement that no controlled comparison had been run was true when written and stale from 14 August; the library revision published with this note carries the measurement and both audits. If we have something wrong here, we correct it in writing on this page.