The agent setup that most often produced working code in the study did so in 57% of tasks. In only 11.8% of tasks did its code both work and pass the security test. Among its working solutions, 79.3% still failed a security test. When agents were not prompted about security, no setup produced code that both worked and passed the security test in more than 12.9% of tasks.
Those are the headline figures. To understand what they mean, we need to look at how the benchmark was built.
The study is Is Vibe Coding Safe? Benchmarking Vulnerability of Agent-Generated Code in Real-World Tasks (Zhao, Wang, Zhang, Luo, Li and Li, ICML 2026, arXiv:2512.03262).
What the benchmark did
The authors built 186 tasks from large open-source repositories on GitHub. Each began with a real case: a human implemented a feature, the implementation proved vulnerable, and the vulnerability was later fixed. Together, the tasks cover 79 weakness categories from MITRE’s Common Weakness Enumeration.
For each task, the agent receives the codebase and a description of the feature, “without implementation details or security guidance”. It is then asked to implement the feature. Its patch is run against two sets of human-written tests. One checks whether the feature works; the other checks for the vulnerability removed by the project’s own fix.
The authors tested three agent frameworks (SWE-agent, OpenHands and Claude Code) with four models: Claude 4 Sonnet, Kimi K2, Gemini 2.5 Pro and Gemini 3 Pro. That gave twelve setups, each with one attempt per task.
What it found
- The setup that most often produced working code, SWE-agent with Claude 4 Sonnet, did so in 57.0% of tasks. In 11.8%, the code also passed the security test. As the authors put it, “79.3% of these functionally correct solutions are insecure”.
- OpenHands with the same model showed a similar pattern: 77.8% of its working solutions were insecure.
- Claude Code with Claude 4 Sonnet produced working code in 41.4% of tasks. In 5.9%, the code also passed the security test.
- When agents were not prompted about security, the highest share of tasks solved both correctly and securely was 12.9%, for SWE-agent with Gemini 3 Pro, which produced working code in 37.1% of tasks.
- Across the twelve setups, that share averaged 8.4% (our arithmetic from Table 3); the authors describe the average as “only around 10%”.
- In every one of the twelve setups, most working solutions failed a security test. The proportion ranged from 59% to 86%, by our arithmetic from the paper’s Table 3.
Telling the agent about security helped a little
The authors tried two prompts on the setup that most often produced working code. The first gives the agent the full list of CWE categories covered by the benchmark, with descriptions. Before writing code, the agent identifies which weaknesses the task may involve. The second prompt names the weakness type or types the task targets and asks the agent to avoid vulnerable implementations.
With these prompts, the share of tasks yielding code that both worked and passed the security test rose from 11.8% to 14.5% and 15.1%, respectively. But the share yielding working code fell from 57.0% to 50.0% and 55.4%. In the authors’ words, the improvement “does not substantially mitigate the security problem”.
In our reading, the second prompt gives the agent information a user does not yet have. The authors note that human experts can identify potential security risks before implementation; the first prompt follows that approach. The second names the weakness type the task was built to test. Outside a benchmark, that information becomes available only once the flaw has been found.
The test existed because the flaw had already been found
This is the part of the design we would underline. Each task comes with a security test added by the project’s developers in the commit that fixed the original vulnerability. Someone had already found the flaw, and the project’s developers had written a test for it. That history is what allows the benchmark to check for the vulnerability. The authors then checked every security test by hand and revised those tied too closely to one implementation.
Your agent’s new feature has no such test. Your suite tells you what the benchmark’s functional tests told the authors: that the feature works. Yet most working solutions in the study still failed a security test.
Finding a flaw nobody has written a test for yet is the work of an adversarial review. Our explanation of adversarial code review sets out what that work covers and where its limits lie.
Different models avoid different flaws
The authors report that “agent frameworks and LLMs are good at avoiding different vulnerabilities”. Kimi K2 handled cryptographic failures better, while Gemini 3 Pro was better at enforcing access control.
That result concerns writing code. It does not establish whether a second model would catch flaws in the first model’s work. Nor is a second model an independent check by default: when two models are wrong, they are often wrong in the same way (Kim et al., ICML 2025). See The attacker can run any model. How many do you run?.
It is one reason we do not treat agreement between models as proof: the claim that it is proof is one of the five claims we refuse.
A question the study leaves open
Every task began with a vulnerability a human had introduced into a real project. In most working solutions, the agents failed the security test for that vulnerability. But the study does not ask whether the models had encountered those projects’ histories, including the vulnerable versions, during training. See The Illusion of the Score.
It therefore cannot distinguish between reproducing a flaw encountered during training and arriving independently at the same mistake. Our research paper, Agentic Stemmatics, examines how to distinguish those possibilities.
What it does not show
- Base rates. The 186 tasks were chosen because each is security-sensitive and a human once got it wrong. The rates describe that kind of feature, not agent-written code in general.
- Newer models. The study tested four models, with one attempt per task for each setup. Other models, and later versions, may score differently.
- Review. The study measures an agent writing code without human review. It does not tell us what a team, a tool or a reviewer would catch afterwards. See A Clean Benchmark Is Not a Clean Codebase and The Speed Fades. The Complexity Doesn’t..
The result is limited, but clear. In these security-sensitive tasks, most working solutions still failed a security test. A test suite that checked only whether the feature worked would not reveal that distinction.