Code can work perfectly and still let someone steal funds, read private data or take control.
CounterProof is an independent team that tries to break your code before someone else does. For every problem we report, we show the evidence: where it is in the code, how we checked it, and what is still uncertain. You get a signed report you can share with customers, insurers, investors and regulators.
A model checking its own work is not an independent review.
Get an independent review →Nobody buys a code review. They buy what it unlocks.
You're not reading this because you woke up wanting a security assessment. Something is asking you for evidence.
Three of these four never accept your own word about your own code; independence is the property they are buying, and it is the one property no team can supply for its own work. The fourth, the regulator, will often take your self-assessment, and then hold you to every record behind it. Either way the evidence has to hold. That is what we produce.
Where review sits in your vulnerability management, stage by stage →
Finding got cheap. Proving got hard.
Each engagement produces a CounterProof Assessment under a persistent engagement identifier (CPR-YYYY-NNN) that your maintainers, your acquirer, or your fix commits can cite. Every finding is either confirmed against your source (exact file and line, reproduction path, impact classification) or explicitly graded as plausible-only. Nothing padded, nothing scanner-generated, nothing you can't act on. The record also states what we checked and did not find: a recorded negative is a deliverable here, not a failed engagement.
Every finding names the evidence rung it actually reached: source trace, compile-proof (the build made to fail, then pass, exactly as the finding predicts), test, or live reproduction, and for any probabilistic claim a rate with an interval and the per-trial logs kept, and never implies a higher one. Every path:line citation is machine-resolved against the exact revision we reviewed before the report leaves our hands. You can check our work. That is deliberate.
We sit where a finding becomes a decision someone will later be asked to justify: adjudication, closure, and the signed record; not in the scanner's race for candidates. Finding is the by-product of an engagement; the record is the product.
When a finding does not survive scrutiny (yours, your maintainers', or our own re-examination), we retract it in writing, with the reason, under the same engagement identifier. We do this unprompted; the counterparty does not have to ask. The withdrawal log is not an admission; it is the mechanism.
Specimen: the shape of a findingIllustrative specimen, not from a client engagement; identifiers are placeholders.
“The witnesses are now free. The critical edition is not.”
The market has started to compete on how many findings a tool can generate, but a finding is a witness, not a verdict. A file:line citation is something a reader can open and check in a minute; a bare “Critical” is not. In our records, severity is an output of the evidence grade, never an input to the marketing. The witnesses are now free. The critical edition is not. A finding is not a verdict →
Four steps. You know the scope before you pay, and you keep the record after.
A first engagement doesn't have to be everything you ship. A small, well-chosen scope (one module, one trust boundary) produces the same defensible record at a size you can evaluate, and it can be the first piece of Part II(3) evidence a manufacturer files.
Fluency is not a sign of quality. It is the failure mode.
When a person writes a line of code, they doubt it. They know they were guessing, and that doubt is a safety feature, because it is what makes them go and test the thing.
A model writes the same line fluently, in precisely the voice it uses for the lines that are correct. To the person reading, fluency reads as competence. A junior engineer hands you something visibly rough and you check it. A model hands you two hundred polished lines with a confident reason for each, and you don't.
The trap isn't that models are wrong more often than people are. It's that their wrong answers arrive wearing the same voice as their right ones, indistinguishable from the outside, because they are indistinguishable from the inside.
Then the same process writes the test that certifies the code, and the comment that explains it. All three agree, because all three came from one place. That agreement reads like three independent confirmations. It is one.
“Their wrong answers arrive wearing the same voice as their right ones.”
So the useful question isn't whether a model can write good code; it can, with a method around it. It's what your review process is actually measuring when everything it inspects came from the same source: the code, the test, the explanation, and increasingly the tool you built to check them. Done properly, the output is good. Done badly, it reads exactly the same.
A review finds what its reviewer can find.
Run one model over a codebase and a short report tells you something narrow: that model, with those prompts and tools, on that day, surfaced these. What it did not surface is not in the report, and by construction cannot be.
Run the same model on the same code again and expect a different set. LLM code review behaves more like fuzzing than like a static analyser: expect different findings from an exact rerun, from a small change to the prompt, and from a new model generation; and, like fuzzing, it is never finished. (Thomas Dullien, “An age of experimentation”, BlueHat Asia 2026.)
How much that matters depends on how far the blind spots of different reviewers overlap; for code review, nobody has measured it. Research on correlated errors, in model populations that include different vendors, found substantial overlap on multiple-choice benchmarks and a résumé-screening task; whether that carries over to finding vulnerabilities is not established. What it does show is that different vendors are not automatically independent of each other.
Several things can bound what a single pass missed: a proof, a test suite, a fuzzer, a specialist, seeded defects, or a reviewer that differs from the first. What cannot bound it is reading the same pass more confidently.
We have not shown that multi-reviewer review catches more real defects; the experiment we designed to test that was retired before it ran, and we say so publicly. The narrower point is the one we stand behind: a clean pass from one reviewer is a scoped result, and reading it as a cleanliness certificate is an inference that result does not support. The full argument →
We have written the long-form argument down, including where our own method's limits are and what we have not demonstrated: a plain-language introduction, and the full paper behind it, published as a preprint under doi:10.5281/zenodo.22030516 (CC BY 4.0). It is not peer-reviewed, and it says so.
You could run the same models. You can't be independent.
Running multiple AI models over your own code replicates our tooling and loses the property that matters. Your team chooses what the reviewers see, frames the questions, and judges the answers, and every one of those choices carries your assumptions straight back into the review. That is not a discipline failure; it is structural. The author of a system cannot be its own adjudicator.
So even a flawless internal review is still your own word about your own code. What they are buying is a third party willing to sign. We are not a notified body and we do not certify conformity; we produce the independent evidence that supports your assessment and is structured to be re-examined by anyone else's.
The method doesn't care who (or what) wrote your code. Independence is missing from human-written code just as often. Machine-written code only makes the gap impossible to ignore.
You would find this anyway. Better you hear it from us.
CounterProof is part of a group that builds payment and digital-asset custody infrastructure. That is where this method was forged, and it means that if you build in those markets, our affiliate may be adjacent to you, or competing with you.
So: before any engagement, we tell you exactly what the group builds and where it operates, in writing. You decide whether that is acceptable, and you decide before you have paid us anything. If it is not acceptable, that is a legitimate answer and we would rather hear it at the start.
What our independence claim does and does not cover, precisely: we do not issue an assessment on code authored by CounterProof or by any company in our group. That boundary is structural. It is a statement about whose code we review, not a claim to have no commercial interests anywhere near your sector. Ours are named above, and you decide whether they are acceptable before you have paid us anything.
The claims we refuse to make on your behalf are listed separately: What we will not claim →
Three people, named, on the record.
A signed assessment means someone's name is on it. These are the people, and which name carries what.
We are a family-run practice (father, son, and daughter), and the review team is two people. Both are facts a diligence team turns up on its own, so we would rather say them here: a two-person review team takes fewer engagements than a firm, declines everything outside its domain, and cannot hide a weak pass behind a brand. That is the trade, and it is the reason the method is written down and mechanised rather than carried in someone's head.
What we are reading, and what we are checking.
Other vendors' models helped write, select and judge Kolibri-1's training data. The EU summary lists seven names; the technical report explains the different roles. These disclosures show how a model built for sovereign and specialised deployment can draw on other vendors' models during training. They do not establish which behaviours, if any, were inherited.
The Regulation never mentions crypto wallets. It applies to products with digital elements. From December 2027, it also requires effective and regular tests and reviews of the security of the product. In our reading, that duty bears on whether a maker finds a flaw before an attacker does. This note explains the scope, the reporting deadlines and the role of an independent review.
Of twelve agent setups, the one that most often produced working code did so in 57% of tasks; in only 11.8% did its code both work and pass the security test. Telling the agent about the risk helped a little. The study could test for each flaw because it had already been found and fixed, and the fix came with a security test. Your agent's new code has no such test yet.
Talk to us before the deadline does.
Email us what you need reviewed, what decision the assessment must support, and when you need it. Elenora Soons, our account manager, will arrange a call with the review team.