CounterProof

Adversarial code review: what the term leaves out

The phrase now covers four practices, from a critic prompt to an attacker-minded human house. Each supplies the attacker's framing; none by itself supplies a signed record from reviewers independent of your author, graded by the evidence each finding reached. What we deliver, what a second model cannot, and what we tell you before you buy.

Adversarial code review is reading source with an attacker’s question in mind (how could this be made to do something it should not) rather than a maintainer’s question of whether the code is correct and tidy. What we deliver under that name is a signed finding record: every confirmed finding graded by the evidence it actually reached, every surface examined and found sound written down with the method that examined it, produced by reviewers independent of your code’s author, adjudicated by a named person your acquirer, insurer or customer can cite, and withdrawn in writing if it does not survive. This page says what that adds to the other things the phrase now means, and what it costs.

What the phrase means today

In our reading of the market (a search in September 2026, not a census) four practices carry the name. Each does something real. The first two are instruments on our own bench.

The critic prompt. One agent writes a change; a second, in fresh context and often from another model family, reads the diff through hostile lenses. This removes self-grading, which matters where it applies: measured across eight judge models, self-preference was large for one, moderate for two, near zero for two and slightly reversed for three (self-preference bias, arXiv:2410.21819). As a practice it supplies the attacker’s framing; it does not by itself provide a retained record, a measure of how correlated the two readings were, or a signature.

The two-reviewer orchestrator. Two model families review the same change independently (sometimes blind to each other) then cross-critique, and one of them synthesises the result, often implementing the fix. The blind phase is right. As a shape it leaves the reviewing model as the judge of its own findings and their repairman, and it does not by itself produce ground truth or a correlation measure.

The platform ensemble. A review product runs several models inside its own pipeline and returns pull-request comments with severities; it carries the name mostly by adjacency, the phrase’s self-describers are overwhelmingly the first two practices. Model plurality here is a vendor design decision; as sold, the shape does not expose which model said what, on what evidence, or how often they agree when wrong. It supplies volume and integration. A comment thread is not an audit.

The attacker-minded human house. Readers with offensive backgrounds review the source the way an attacker would, whether as a firm or as a crowd. This is where an attack nobody framed gets invented, and it is the practice ours is built to feed rather than replace. As a shape it does not by itself separate its evidence from its verdict, and its reviewers’ independence from each other is unmeasured, as everyone’s is.

Three properties, one word

Three properties get bundled into adversarial: the task framing (the attacker’s question), the error independence between reviewers, and the organisational independence of whoever signs.

Task framingError independenceOrganisational independence
Critic promptyesnot by itself, one brief, one judgeno
Two-reviewer orchestratoryesa blind phase, unmeasured; the reviewer arbitratesno
Platform ensembleyesvendor-internal, not exposedno, a tool cannot sign
Human houseyesone teamyes
CounterProofyesbounded and recorded, not measured: separate model lineages, the brief on file, correlated or failed seats not countedyes: named signatory independent of the code’s author; sector interests disclosed in writing

A second model, chosen by you, pointed at what you show it, judged by you, is not an adversary. It is a second draw from a distribution you are already standing inside. Three gaps follow, and none is a model-quality problem.

Gap one: a different vendor is not an independent witness

Two models from different companies have different training data, so agreement is corroboration and silence is reassurance; that is the reasoning. One study across more than 350 models found that when two models both get something wrong they agree on the same wrong answer far more often than chance (a mean of 60% on one leaderboard and 42% on the other) with pairs from different providers included throughout; its cross-provider statement is an inference from the study’s regression model, not a cross-provider-only analysis (Kim et al., ICML 2025, arXiv:2506.07962). It measured benchmarks and résumé screening, not vulnerability review, so it gives nobody the overlap figure for code; what it establishes is that different vendors are not automatically independent. That this should change how you read a second reviewer’s agreement on your code is our inference. Textual scholars made the same point a rule five centuries ago: a copy shown to descend from another surviving copy is struck from the reckoning, not because it is wrong but because it adds no testimony.

The correlation channel people miss is not the vendor. It is the brief: prompt, tools, evidence bundle and role framing correlate reviewers across vendors, and two models handed the same curated packet share whatever the packet leaves out. We treat the brief as part of the apparatus and file it with the review. The attacker can run any model works through what this does and does not license.

Gap two: nobody decides when an opinion should not be asked for

Some claims have an oracle, something that answers them without anyone’s opinion. Does this path execute? Run it. Does the specification permit this value? Read it. Where an oracle exists, a test that ran outranks a model that reasoned, and a person who reasoned too.

Asked whether a regex matches a path, a reviewer agent produces a fluent opinion about whether the regex matches the path. Routing that question to an execution instead is an engineering choice, and in our reading most review tooling has not made it, because the category is bought by findings per pull request, and a finding that turns out to be a two-minute script is not a finding. We know the difficulty from the inside: we built a check to enforce this routing and shipped it with its self-test defined and never called. What settles a finding is the autopsy.

The claims with no oracle (how severe, is this trust assumption acceptable, would an attacker bother) are judgement, and no volume of model output converts judgement into fact.

Gap three: nobody has signed

An in-house agent run (however many models, however isolated their contexts) produces your own word about your own code, with nobody outside your organisation accountable for a sentence of it. In our experience that is the property an acquirer’s counsel, a cyber-insurer’s underwriter and an enterprise customer’s security team ask for first, because it is the one your team cannot supply about its own work.

Under the Cyber Resilience Act most manufacturers may self-assess and are then held to every record behind it: reports of the tests carried out in the technical documentation (Annex VII(6)), evidencing the testing duty (Annex I Part II(3)), retained ten years or the support period if longer (Article 13(13)). Reporting for actively exploited vulnerabilities began 11 September 2026; the Regulation applies in full from 11 December 2027. An external assessment supports the file you assemble; it is not the file, and we are not a notified body. The standards slipped, the deadline did not covers the timeline.

What you get from us

None of the following is a secret. A well-run team could adopt any one of them. What you are buying is that all of them are done, written down, and signed, as a record you can hand to someone who was not in the room.

  1. Oracle first. A decidable claim is routed to execution, compilation or the specification before any opinion is commissioned on it. Our review-pack tooling refuses to build a brief without its evidence bundle, and the brief states what that bundle cannot show, so a seat’s silence on something it was never given is read as silence, not as absence.
  2. Seats are lineages, and the brief is on file. Reviewers are separate model families, seats with file access building their own evidence from source, seats without it reading only what the brief carries. The brief, tools and bundle each received are part of the record. A seat that failed is recorded absent, never as agreement; a seat that turns out to share context with another is flagged and not counted, including the time one of ours turned out to have our own findings document in its context.
  3. Blind returns are captured before adjudication exists. What each seat said, verbatim, with its envelope (model, harness, date, files opened) is written down before anyone compares them.
  4. A tally never establishes anything. Findings do not close by vote. Where reviewers disagree about what the code does, the source decides; where a rejection rests on “uncertain”, the claim goes to a retained queue with an owner, an expiry and a re-open criterion, and is re-exposed to a fresh lineage. Confirmation runs up the evidence rung: source trace, compile-proof, test, live reproduction, and for any probabilistic claim a rate with an interval and the per-trial logs kept; anything short of a rung is plausible-only, the other disposition rather than a weaker rung. A surface examined and found sound is recorded with the method that examined it and the control that showed the method can find what it looked for.
  5. A named person signs, and withdraws in writing. Under an identifier your maintainers and your acquirer can cite, including when the finding was ours. We publish the cases where our own instruments failed, because a practice that cannot show you where it was wrong has no standing where it says it is right.

Polling several models and taking the majority is an ensemble; as sold, it exposes none of this to the buyer. When we searched in September 2026 we did not find any practice offering this bundle as a protocol a buyer can inspect. That is a statement about shape, not about catch rate: what we tell you before you buy, below, is about that.

Where this sits in your vulnerability management

Finding is not where we earn our fee. Scanners, fuzzers, review platforms and your own engineers already produce candidates faster than anyone can adjudicate them, and our engagements find things too; that is the by-product. What we are paid for sits where the lifecycle breaks: adjudication: a candidate is routed to its oracle first, then graded by the rung of evidence it reached, so severity follows proof rather than a model’s confidence; closure: disagreement is settled by the source, not by vote, and “uncertain” is held in a retained queue with an owner instead of quietly becoming “accepted”; and assurance; the outcome is a record a named person outside your organisation signed and can withdraw, the part of a technical file, a diligence pack or an underwriting submission you cannot produce about your own work. Place us in identification and we compete with your scanner on volume, a race we do not run. Place us where a finding becomes a decision someone will later be asked to justify, and the record is the product. A finding is not a verdict covers the first stage and What settles a finding the settling half of the second; the retained queue is described above. Vulnerability management under the Cyber Resilience Act walks every stage with the CRA’s duty at each.

What we tell you before you buy

These protect you, so they are on the page rather than in the contract.

  • We sell a record, not a catch rate. We have not shown that a panel finds more real defects than one good reviewer, and we do not claim it; we measured our full apparatus against one cheap pass, pre-registered and blind-scored, and it won neither target outright: it lost one, and on the other it met its own precision trigger only by deleting the findings the auditor had ruled real. We measured our own method, and it did not earn its cost. The case report we work from says a simpler doctrine (one independent outside review on anything you will act on, disputes settled by reading the source, no unmeasured number shipped) fits the same evidence; if that is all you take from this page, take it. What we sell is the record of having done it, under a signature. Independence between our seats is bounded by procedure and recorded, not measured for code review; nobody has measured it. Every finding in the record reached a named rung or is marked as not having done so.
  • Where the author is in the loop. Our rule is that whoever raised a disputed claim does not adjudicate it alone. Our reviewing bench is two people, which means fewer engagements, declining work outside our domain, and that this rule has no mechanism enforcing it yet; our own case report names that absence as its largest unfixed weakness. In practice a disputed claim goes to the source and to a fresh lineage before it is closed, the adjudication is recorded, and the signatory is named, so you can see which findings were closed that way and weigh them yourself.
  • No method reaches a defect no witness raised. Collation selects among what was emitted. Tests, fuzzing, proofs and a specialist’s reading are separate routes to the same place, and a serious programme uses them too.
  • We are independent of your code’s author, not disinterested in every sector. CounterProof is part of a group that builds payment and digital-asset custody infrastructure. If you build in those markets, our affiliate may be adjacent to you or competing with you. We tell you in writing what the group builds and where it operates before any engagement and before any money moves, and we issue no assessment on code authored by CounterProof or any company in our group.

The long-form argument, with its own limits, is our working paper, a preprint, not peer-reviewed: research.counterproof.io (doi:10.5281/zenodo.22030516).

Questions we get

Is adversarial code review the same as penetration testing? No. Penetration testing exercises a running system; this reads source at a pinned revision and establishes findings against the code. Both are attacker-minded, they answer different questions, and a serious programme uses both.

Can I not just run two AI models over my own code? You can, and a second reviewer may surface candidates the first did not. What it cannot produce is organisational independence: your team chooses what the reviewers see, frames the questions, and judges the answers, and nobody outside it has signed.

A review platform already runs several models. Isn’t that a panel? No. An ensemble whose models the vendor chooses, whose individual returns you cannot see, and whose disagreements resolve inside the pipeline is one instrument with several parts. A panel is several instruments whose returns are recorded separately and whose disagreements are settled by the source.

Does adversarial code review only apply to AI-generated code? No. The method does not care who (or what) wrote the code. Machine-written code makes the gap harder to ignore, because volume rises and fluency hides the guesswork.

How is this different from an AI code review tool? Those tools generate findings, and we run tools in that class as instruments. What a tool cannot do is be accountable or sign.


Inferences marked as such above, and here: that a second reviewer’s agreement on code should be discounted (Kim et al. measured other domains); that the brief correlates reviewers; that the four practices are the market’s shape and that review tooling competes on findings per pull request (our reading, one search, September 2026); that organisational independence is what acquirers, insurers and enterprise buyers ask for first (our experience, not a survey). If we have something wrong here, we will correct it in writing on this page.

← All notes