CounterProof

Your AI wrote the code. Who checked it?

Code can work perfectly and still let someone steal funds, read private data or take control.

CounterProof is an independent team that tries to break your code before someone else does. For every problem we report, we show the evidence: where it is in the code, how we checked it, and what is still uncertain. You get a signed report you can share with customers, insurers, investors and regulators.

A model checking its own work is not an independent review.

Get an independent review →
01

Why you're here

Nobody buys a code review. They buy what it unlocks.

You're not reading this because you woke up wanting a security assessment. Something is asking you for evidence.

A regulator
The EU Cyber Resilience Act requires you to “apply effective and regular tests and reviews of the security” of your product (Annex I Part II(3)), and your technical documentation must contain “reports of the tests carried out to verify conformity” (Annex VII(6)), kept for at least ten years after placing on the market, or the support period if longer (Article 13(13)). Article 14 notification of actively exploited vulnerabilities and severe incidents starts 11 September 2026; the Regulation applies in full from 11 December 2027. Most products may self-assess, which is not relief but exposure: the burden of producing records that stand up falls entirely on you.
An acquirer or investor
Technical diligence is where AI-assisted codebases now get discounted. An independent assessment on the table changes that conversation before it starts.
A cyber-insurer
Underwriters increasingly price the gap between “we test internally” and “an independent party tested and signed.”
An enterprise customer
Their security questionnaire has no checkbox for “our model reviewed its own output.”

Three of these four never accept your own word about your own code; independence is the property they are buying, and it is the one property no team can supply for its own work. The fourth, the regulator, will often take your self-assessment, and then hold you to every record behind it. Either way the evidence has to hold. That is what we produce.

Where review sits in your vulnerability management, stage by stage →

02

What you get

Finding got cheap. Proving got hard.

Each engagement produces a CounterProof Assessment under a persistent engagement identifier (CPR-YYYY-NNN) that your maintainers, your acquirer, or your fix commits can cite. Every finding is either confirmed against your source (exact file and line, reproduction path, impact classification) or explicitly graded as plausible-only. Nothing padded, nothing scanner-generated, nothing you can't act on. The record also states what we checked and did not find: a recorded negative is a deliverable here, not a failed engagement.

Every finding names the evidence rung it actually reached: source trace, compile-proof (the build made to fail, then pass, exactly as the finding predicts), test, or live reproduction, and for any probabilistic claim a rate with an interval and the per-trial logs kept, and never implies a higher one. Every path:line citation is machine-resolved against the exact revision we reviewed before the report leaves our hands. You can check our work. That is deliberate.

We sit where a finding becomes a decision someone will later be asked to justify: adjudication, closure, and the signed record; not in the scanner's race for candidates. Finding is the by-product of an engagement; the record is the product.

When a finding does not survive scrutiny (yours, your maintainers', or our own re-examination), we retract it in writing, with the reason, under the same engagement identifier. We do this unprompted; the counterparty does not have to ask. The withdrawal log is not an admission; it is the mechanism.

Specimen: the shape of a finding
CPR-2026-0NN · F-07 — duplicate credit on settlement retry severity: High (reachable in default build) evidence rung: test — failing test, 2/2 runs, 3 clean negative controls source: gateway/src/settle.rs:412 @ rev a1b2c3d fix verified: rev e5f6a7b — finding closed in writing

Illustrative specimen, not from a client engagement; identifiers are placeholders.

“The witnesses are now free. The critical edition is not.”

To an authority
Structured to be included in your technical documentation as an Annex VII(6) test report, evidencing the Part II(3) testing duty. It supports the file you assemble; it is not the file, and it does not verify conformity across Annex I.
To a diligence team
Evidence-graded, reproducible, scoped, with our name on the risk.
To your own engineers
Short enough to read, precise enough to fix. We can stay through the patch and independently verify each fix against the finding it targets, recording what we verified and at which revision, so you end with a record of what was checked, not a list of notes.

The market has started to compete on how many findings a tool can generate, but a finding is a witness, not a verdict. A file:line citation is something a reader can open and check in a minute; a bare “Critical” is not. In our records, severity is an output of the evidence grade, never an input to the marketing. The witnesses are now free. The critical edition is not. A finding is not a verdict →


And this is how an engagement runs

Four steps. You know the scope before you pay, and you keep the record after.

1. Scope
A call with the review team, the people who will do the work. We agree what is in and out of scope, pin the exact revision under review, and put the group disclosure in front of you in writing, before any money moves. If the work is outside our domain, we say so and decline.
2. Review
The review team works the pinned revision. Every candidate finding is adjudicated against the source (confirmed, graded plausible-only, or discarded), and the checks that came back clean are recorded alongside the ones that didn't.
3. Report
You receive the assessment under its engagement identifier, with a walkthrough for your engineers. Findings, evidence rungs, reproduction paths, recorded negatives: short enough to read, precise enough to act on.
4. The record
We can stay through your patches and verify each fix against the finding it targets, at the revision that closed it. The identifier, the withdrawal log, and the closure trail remain citable (to your customer's security team, your acquirer, or your authority) under the retention terms stated in the engagement agreement.

A first engagement doesn't have to be everything you ship. A small, well-chosen scope (one module, one trust boundary) produces the same defensible record at a size you can evaluate, and it can be the first piece of Part II(3) evidence a manufacturer files.

03

What actually goes wrong

Fluency is not a sign of quality. It is the failure mode.

When a person writes a line of code, they doubt it. They know they were guessing, and that doubt is a safety feature, because it is what makes them go and test the thing.

A model writes the same line fluently, in precisely the voice it uses for the lines that are correct. To the person reading, fluency reads as competence. A junior engineer hands you something visibly rough and you check it. A model hands you two hundred polished lines with a confident reason for each, and you don't.

The trap isn't that models are wrong more often than people are. It's that their wrong answers arrive wearing the same voice as their right ones, indistinguishable from the outside, because they are indistinguishable from the inside.

Then the same process writes the test that certifies the code, and the comment that explains it. All three agree, because all three came from one place. That agreement reads like three independent confirmations. It is one.

“Their wrong answers arrive wearing the same voice as their right ones.”

So the useful question isn't whether a model can write good code; it can, with a method around it. It's what your review process is actually measuring when everything it inspects came from the same source: the code, the test, the explanation, and increasingly the tool you built to check them. Done properly, the output is good. Done badly, it reads exactly the same.


And the attacker is not limited to your reviewer

A review finds what its reviewer can find.

Run one model over a codebase and a short report tells you something narrow: that model, with those prompts and tools, on that day, surfaced these. What it did not surface is not in the report, and by construction cannot be.

Run the same model on the same code again and expect a different set. LLM code review behaves more like fuzzing than like a static analyser: expect different findings from an exact rerun, from a small change to the prompt, and from a new model generation; and, like fuzzing, it is never finished. (Thomas Dullien, “An age of experimentation”, BlueHat Asia 2026.)

How much that matters depends on how far the blind spots of different reviewers overlap; for code review, nobody has measured it. Research on correlated errors, in model populations that include different vendors, found substantial overlap on multiple-choice benchmarks and a résumé-screening task; whether that carries over to finding vulnerabilities is not established. What it does show is that different vendors are not automatically independent of each other.

Several things can bound what a single pass missed: a proof, a test suite, a fuzzer, a specialist, seeded defects, or a reviewer that differs from the first. What cannot bound it is reading the same pass more confidently.

We have not shown that multi-reviewer review catches more real defects; the experiment we designed to test that was retired before it ran, and we say so publicly. The narrower point is the one we stand behind: a clean pass from one reviewer is a scoped result, and reading it as a cleanliness certificate is an inference that result does not support. The full argument →

We have written the long-form argument down, including where our own method's limits are and what we have not demonstrated: a plain-language introduction, and the full paper behind it, published as a preprint under doi:10.5281/zenodo.22030516 (CC BY 4.0). It is not peer-reviewed, and it says so.

04

Why this can't be done in-house

You could run the same models. You can't be independent.

Running multiple AI models over your own code replicates our tooling and loses the property that matters. Your team chooses what the reviewers see, frames the questions, and judges the answers, and every one of those choices carries your assumptions straight back into the review. That is not a discipline failure; it is structural. The author of a system cannot be its own adjudicator.

So even a flawless internal review is still your own word about your own code. What they are buying is a third party willing to sign. We are not a notified body and we do not certify conformity; we produce the independent evidence that supports your assessment and is structured to be re-examined by anyone else's.

The method doesn't care who (or what) wrote your code. Independence is missing from human-written code just as often. Machine-written code only makes the gap impossible to ignore.

05

The disclosure that comes with that

You would find this anyway. Better you hear it from us.

CounterProof is part of a group that builds payment and digital-asset custody infrastructure. That is where this method was forged, and it means that if you build in those markets, our affiliate may be adjacent to you, or competing with you.

So: before any engagement, we tell you exactly what the group builds and where it operates, in writing. You decide whether that is acceptable, and you decide before you have paid us anything. If it is not acceptable, that is a legitimate answer and we would rather hear it at the start.

What our independence claim does and does not cover, precisely: we do not issue an assessment on code authored by CounterProof or by any company in our group. That boundary is structural. It is a statement about whose code we review, not a claim to have no commercial interests anywhere near your sector. Ours are named above, and you decide whether they are acceptable before you have paid us anything.

The claims we refuse to make on your behalf are listed separately: What we will not claim →

06

Who we are

Three people, named, on the record.

A signed assessment means someone's name is on it. These are the people, and which name carries what.

Vincent Soons, CEO
Runs security research at CounterProof Research, focused on Bitcoin infrastructure: L2 protocols, custody systems, and the code that moves real money. His work is backed by runnable test harnesses. Evidence levels are stated per claim, and anything he reports as reproduced has been reproduced deterministically against unmodified code. Background: upstream contributor to Fedimint, the federated Bitcoin custody protocol; public contributions, checkable by anyone.
Luuk Soons, CSO
Active in Bitcoin and sound-money infrastructure since 2013. Owns the method, the regulatory posture, and every disclosure decision, including the one above.
Elenora Soons, Account Manager
Account manager, and your contact for information requests: the person who answers when you write to us, scopes the conversation, and keeps an engagement's paperwork moving. She does not author or sign assessments; the two names above carry those. info@counterproof.io · +356 7741 7951

We are a family-run practice (father, son, and daughter), and the review team is two people. Both are facts a diligence team turns up on its own, so we would rather say them here: a two-person review team takes fewer engagements than a firm, declines everything outside its domain, and cannot hide a weak pass behind a brand. That is the trade, and it is the reason the method is written down and mechanised rather than carried in someone's head.

07

Notes

What we are reading, and what we are checking.

All notes →

08

Talk to us before the deadline does.

Email us what you need reviewed, what decision the assessment must support, and when you need it. Elenora Soons, our account manager, will arrange a call with the review team.

info@counterproof.io

+356 7741 7951

Follow us on LinkedIn Follow us on X