What we are reading in the Cyber Resilience Act, and what we are finding in how machine-generated work gets evaluated. Dated, sourced, and corrected in writing when we get something wrong.
Other vendors' models helped write, select and judge Kolibri-1's training data. The EU summary lists seven names; the technical report explains the different roles. These disclosures show how a model built for sovereign and specialised deployment can draw on other vendors' models during training. They do not establish which behaviours, if any, were inherited.
The Regulation never mentions crypto wallets. It applies to products with digital elements. From December 2027, it also requires effective and regular tests and reviews of the security of the product. In our reading, that duty bears on whether a maker finds a flaw before an attacker does. This note explains the scope, the reporting deadlines and the role of an independent review.
Of twelve agent setups, the one that most often produced working code did so in 57% of tasks; in only 11.8% did its code both work and pass the security test. Telling the agent about the risk helped a little. The study could test for each flaw because it had already been found and fixed, and the fix came with a security test. Your agent's new code has no such test yet.
AI code review is an experimental practice. Dullien's BlueHat Asia talk explains the measurement problem; our work asks what evidence supports each finding. Where the approaches meet, and what we have yet to establish.
TypeSafe AI's Jev is a new kind of model: it never writes text, only returns typed decisions with probabilities. Its benchmark grades it against the average of two other AI models. We opened the cases TypeSafe publishes and counted. In 8 of the 19 cases where both graders answered, the graders disagreed with each other. Every one of the four cases TypeSafe publishes under the label 'all three miss the reference' is one where the reference was itself contested. A plain-language note on what that means for how we judge AI.
Finding got cheap; the rest of the lifecycle did not. A stage-by-stage walk (identify, assess, decide, remediate, report, retain) with what the EU Cyber Resilience Act asks at each, and where an independent adversarial review sits: in the stages after the first.
Two pre-registered, blind-scored comparisons of our full multi-model review against a single cheap pass, with published professional audits as the answer key. The full review lost one target outright; on the other it met its own precision trigger only by deleting the findings the auditor had ruled real (0–1 of 4, depending on the scorer). What we changed, what we re-measured, and the holes a later audit found in the measurement itself.
Four quite different practices now go by this name. All of them look at your code the way an attacker would. None of them, on its own, gives you both a signed report from reviewers outside your organisation and a clear statement of how far each finding has been proven. Below: what we deliver, what a second AI model cannot, and the limits we state before you buy.
A finding is a witness. This note is about the next question: what turns a witness into something you can act on, and what happens when the thing being reviewed is a tool we built ourselves.
AI vulnerability scanners post strong numbers on published benchmarks. Two independent studies (one built specifically to test at repository scale, one that watched professional developers use a tool on their own code) found the strong number does not travel. Here is where it breaks.
AI can now generate thousands of vulnerability findings in an afternoon. A finding is a witness, not a verdict, and the one field nobody can check is the one everybody inflates. Here is how to read the flood without drowning in it.
Three independent studies, three different methods (a targeted vulnerability study, a controlled human-vs-model comparison, and a year-long study of real repositories) converge on the same finding: AI-generated code carries a measurable quality cost that outlasts the productivity gain it arrived with.
A clean AI review tells you what one reviewer surfaced on one day. What it did not surface is not in the report. How much that matters is a question almost nobody has measured, including us.
Contamination, correlated judges, and leaderboard gaming look like three problems. They are one problem wearing three costumes, and there are four questions that see through all of them.
We shipped a reproduction figure that had no retained artifact behind it. We found it, withdrew it in writing, and replaced it with a narrower claim that is actually measured. This is the withdrawal commitment being exercised, on us.
AI has collapsed the cost of finding vulnerabilities. In Bitcoin's open-source ecosystem that collision arrived this month, and Brussels has set a clock on what happens next.
A draft amendment would push the CRA's vulnerability-handling standards later. Article 14 is unmoved, which leaves manufacturers documenting a duty whose yardstick is still being drafted.