CounterProof is an independent adversarial review practice for modern codebases — machine-written code above all. We deliver the one thing you cannot produce in-house: an independent, signed, evidence-graded security assessment, built for the scrutiny of regulators, acquirers, insurers, and enterprise customers.
Code written by a model and reviewed by the same model has been reviewed by nobody.
Nobody buys a code review. They buy what it unlocks.
You're not reading this because you woke up wanting a security assessment. Something is asking you for evidence.
Three of these four never accept your own word about your own code — independence is the property they are buying, and it is the one property no team can supply for its own work. The fourth, the regulator, will often take your self-assessment — and then hold you to every record behind it. Either way the evidence has to hold. That is what we produce.
The deliverable is the point.
An engagement produces a CounterProof Assessment under a persistent engagement identifier (CPR-YYYY-NNN — what your maintainers, your acquirer, or your fix commits cite). Every finding is either confirmed against your source — exact file and line, reproduction path, impact classification — or explicitly graded as plausible-only. Nothing padded, nothing scanner-generated, nothing you can't act on.
Every finding names the evidence rung it actually reached — source trace, compile-proof, test, or live reproduction — and never implies a higher one. Every path:line citation is machine-resolved against the exact revision we reviewed before the report leaves our hands. You can check our work. That is deliberate.
Fluency is not a sign of quality. It is the failure mode.
When a person writes a line of code, they doubt it. They know they were guessing — and that doubt is a safety feature, because it is what makes them go and test the thing.
A model writes the same line fluently, in precisely the voice it uses for the lines that are correct. To the person reading, fluency reads as competence. A junior engineer hands you something visibly rough and you check it. A model hands you two hundred polished lines with a confident reason for each, and you don't.
The trap isn't that models are wrong more often than people are. It's that their wrong answers arrive wearing the same voice as their right ones — indistinguishable from the outside, because they are indistinguishable from the inside.
Then the same process writes the test that certifies the code, and the comment that explains it. All three agree, because all three came from one place. That agreement reads like three independent confirmations. It is one.
So the useful question isn't whether a model can write good code — it can, with a method around it. It's what your review process is actually measuring when everything it inspects came from the same source: the code, the test, the explanation, and increasingly the tool you built to check them. Done properly, the output is good. Done badly, it reads exactly the same.
You could run the same models. You can't be independent.
Running multiple AI models over your own code replicates our tooling and loses the property that matters. Your team chooses what the reviewers see, frames the questions, and judges the answers — and every one of those choices carries your assumptions straight back into the review. That is not a discipline failure; it is structural. The author of a system cannot be its own adjudicator.
So even a flawless internal review is still your own word about your own code. What they are buying is a third party willing to sign. We are not a notified body and we do not certify conformity; we produce the independent evidence that supports your assessment and is structured to be re-examined by anyone else's.
The method doesn't care who — or what — wrote your code. Independence is missing from human-written code just as often. Machine-written code only makes the gap impossible to ignore.
You would find this anyway. Better you hear it from us.
CounterProof is part of a group that builds payment and digital-asset custody infrastructure. That is where this method was forged — and it means that if you build in those markets, our affiliate may be adjacent to you, or competing with you.
So: before any engagement, we tell you exactly what the group builds and where it operates, in writing. You decide whether that is acceptable, and you decide before you have paid us anything. If it is not acceptable, that is a legitimate answer and we would rather hear it at the start.
What our independence claim does and does not cover, precisely: we do not issue an assessment on code authored by CounterProof or by any company in our group. That boundary is structural. It is a statement about whose code we review, not a claim to have no commercial interests anywhere near your sector — no review firm with real domain expertise can honestly claim the latter, and we are not going to pretend otherwise.
Two people, named, on the record.
A signed assessment means someone's name is on it. These are the names.
We are father and son, and we are two people. Both are facts a diligence team turns up on its own, so we would rather say them here: a two-person practice takes fewer engagements than a firm, declines everything outside its domain, and cannot hide a weak pass behind a brand. That is the trade, and it is the reason the method is written down and mechanised rather than carried in someone's head.
What we are reading in the Regulation.
AI has collapsed the cost of finding vulnerabilities. In Bitcoin's open-source ecosystem that collision arrived this month — and Brussels has set a clock on what happens next.
A draft amendment would push the CRA's vulnerability-handling standards later. Article 14 is unmoved — which leaves manufacturers documenting a duty whose yardstick is still being drafted.