# An age of experimentation: Thomas Dullien's talk, and where our work meets it

> AI code review is an experimental practice. Dullien's BlueHat Asia talk explains the measurement problem; our work asks what evidence supports each finding. Where the approaches meet, and what we have yet to establish.

Source: https://counterproof.io/notes/age-of-experimentation/ · Published 23 September 2026 · CounterProof is a practice of Clavestra Capital Limited (Malta, C 113987).


Run an AI reviewer twice on the same code, and it may return different findings. Change a few words in its instructions, and the results may change again. How, then, do we know whether we have improved the review process?

Thomas Dullien's BlueHat Asia 2026 talk, *An age of experimentation*, addresses this difficulty. Dullien, of OpenAI's Cyber Security Research Team, argues that working with reasoning language models requires experimental discipline. A change that seems useful needs measurement. What feels like progress may be noise.

Our work approaches a related problem through textual criticism. This note explains where the two approaches meet, the further question our work asks, and what we have yet to measure.

## Learning to use fire

Dullien distinguishes technologies built from understood principles from those discovered before their workings were understood. The internal combustion engine belongs to the first category; fire belongs to the second.

Reasoning language models, he argues, are closer to fire. "We discovered them more than we constructed them." Our understanding remains limited, and improvements come largely through experiment. We are learning what these systems can do while still learning how to explain them.

## When the same input no longer promises the same result

Software engineering relies heavily on predictable behaviour. We change a program, run its tests, and expect the results to tell us whether its behaviour has changed. Language models complicate that expectation.

Dullien places this in a longer history. Rowhammer exposed limits to reliable hardware behaviour. Spectre showed how differences in execution timing could leak information across security boundaries. Language models introduce uncertainty into the meaning and content of the output itself: what he calls the loss of semantic determinism.

Sampling can in theory be reproducible, but practical systems have further sources of randomness, including hardware and concurrency. Dullien also describes models as chaotic: a small prompt change can alter a long sequence of subsequent steps in ways we cannot predict analytically.

His conclusion is practical. We need experiments to establish what a change has done.

## What an improvement takes to measure

In a probabilistic system, Dullien says, "every change is a hypothesis test". Before calling a revision an improvement, we need to understand how much the results vary when nothing has changed.

His sample-size rule makes the difficulty clear: the number of samples needed is proportional to **the square of the ratio of standard deviation to the effect we want to detect**. Standard deviation measures the noise. A small improvement against substantial noise can therefore be expensive to confirm.

The environment can complicate the comparison. In his cloud-performance example, other workloads on a shared host can introduce roughly ±15% variation in runtime. They can also affect several runs together, undermining the assumption that each measurement is independent.

Where possible, he recommends paired experiments: run the two variants side by side on the same virtual machine, exposed to the same noise. He calls this the "tomato farmer protocol", borrowing from fertiliser testing. When individual changes have effects too small to measure economically, he suggests testing five or six suspected improvements together.

Bug detection presents another problem: what counts as a trial? Suppose an agent searches a codebase containing a hundred known bugs. If it skips a file, several bugs may go undiscovered for the same reason. The hundred bugs are therefore not a hundred independent trials. Dullien asks whether the whole run should instead count as one.

Testing bugs separately does not necessarily remove the dependence. Success on several may require the same underlying ability. Treating those outcomes as independent can make confidence intervals appear narrower than they should be.

This is why ordinary continuous integration and delivery processes fit prompt changes poorly. A successful rerun may tell us less than we think. Dullien does not demand a costly statistical demonstration for every small change. His warning is against spending months making adjustments that feel productive without establishing that the system has improved.

## Code review as a continuing search

Dullien compares LLM code review to fuzzing rather than traditional static application security testing. Each run explores possibilities, and most APIs do not let the user fix the random seed. Repeated runs, small prompt changes, and new model generations can each uncover different bugs.

Review therefore has no simple point at which it is finished. He proposes tracking useful discoveries per dollar of compute, expecting the rate to fall as the search proceeds.

His other analogy is a fishery. Bugs enter software and are removed from it at different rates. We can measure effort and catches without knowing the total population. Catching everything may be unrealistic; an unsuccessful search does not tell us the waters are empty.

## Where our work meets his

We reached several related conclusions through textual criticism: the study of how texts are copied, altered, and transmitted. Our paper, [*Agentic Stemmatics*](https://doi.org/10.5281/zenodo.22030516), develops that approach.

- **One run is one draw.** The paper treats model output as testimony. A model [lineage](https://counterproof.io/glossary/#lineage) must be sampled repeatedly; each answer is one draw from it. Dullien's account of reruns finding different bugs reaches the same practical concern. A clean pass describes the result of that review. It does not certify the absence of defects ([*The attacker can run any model*](https://counterproof.io/notes/the-attacker-can-run-any-model/)).
- **Agreement does not establish independence.** Dullien warns that shared success factors can make results across bugs appear more informative than they are. We ask the corresponding question about reviewers: different vendors do not by themselves establish independent errors. His example concerns test items; ours concerns witnesses. In both, a shared cause can reduce what multiple observations establish.
- **Methods have to adapt.** He expects improvements in the software and prompts surrounding a model, the harness, to feed into later model training. Our paper identifies a related risk: once methods and artifacts enter training data, detectors can become targets for optimisation. These are different mechanisms with a common consequence: a method's present usefulness is no guarantee of lasting advantage.
- **Nobody knows the denominator.** His fishery captures a limit our paper also acknowledges. We do not know the full population of defects. We therefore refuse to turn an empty finding list into a claim that nothing is wrong ([*What we will not claim*](https://counterproof.io/refusals/)).
- **Measurement can overturn confidence.** In our comparison of the full review apparatus (our multi-model review pipeline) with a single inexpensive pass, the apparatus did not earn its cost ([*We measured our own method*](https://counterproof.io/notes/we-measured-our-own-method/)). That is the risk Dullien describes: considerable activity can feel like progress until it is measured.

## The further question our work asks

Dullien is chiefly concerned with whether a discovery system has improved. Our emphasis here is what happens after it produces a finding. Is the finding real? What evidence supports it? Who is accountable for the judgement?

The result we seek is a record of how each finding was assessed: what was reproduced by a test against code at a fixed revision, what was supported only by reading, and what was withdrawn. That record lets a reader distinguish kinds of evidence that a single score would conceal.

This meets another observation in the talk. Generating hypotheses and implementing tests have become cheaper; running the experiment can now be the bottleneck. Our own shorthand is: "Finding got cheap. Proving got hard."

Agreement between reviewers adds a further question. Its evidentiary weight depends on the independence of their errors. We have not measured that independence for our reviewers. Obtaining answers from different models does not fill the gap.

Where a test, type check, or proof can settle a finding, it takes precedence over opinions about the finding. That is compatible with Dullien's advice. Measuring how well a system discovers problems and establishing whether a particular problem is real are distinct tasks; both matter.

## What we have yet to establish

We have measured our method on two targets, changed it, measured again, and published the results, including weaknesses a later audit found in the measurement itself.

That work leaves substantial questions open. We have not measured independence between our reviewers' errors. We have not established how their results vary across real workloads. And although we record changes to the method, we have not tested each change in the way Dullien describes.

His talk gives us a useful standard for that work: distinguish the improvement we hope we have made from the improvement the experiment establishes.

## The talk

Thomas Dullien, *An age of experimentation*, BlueHat Asia 2026, 17-18 September, Singapore. [Read the slides (PDF)](https://thomasdullien.github.io/about/slides/An-age-of-experimentation-BlueHat-Asia-2026.pdf).

Dullien has not reviewed this note, and it implies no endorsement of our work. Neither he nor OpenAI is affiliated with CounterProof. This is our reading of a public talk.

---

The full argument, including the limits of our method, is in [*Agentic Stemmatics*](https://doi.org/10.5281/zenodo.22030516), published under CC BY 4.0 and not peer-reviewed.

