CounterProof

TypeSafe Jev, Graded by Consensus: When the Witness Stops Writing

TypeSafe AI's Jev is a new kind of model: it never writes text, only returns typed decisions with probabilities. Its benchmark grades it against the average of two other AI models. We opened the cases TypeSafe publishes and counted. In 8 of the 19 cases where both graders answered, the graders disagreed with each other. Every one of the four cases TypeSafe publishes under the label 'all three miss the reference' is one where the reference was itself contested. A plain-language note on what that means for how we judge AI.

How this note was checked. Before publication it was attacked, refute-first, by eight outside reviewer seats from four model families (Moonshot Kimi, xAI Grok, DeepSeek, Alibaba Qwen), each recomputing every count from the vendor’s files. What they broke and how it was fixed is in §5. The technical companion, with method, full counts and limits, is Field Note 1. Seat records and the hashed data archive are held in our reports store and available on request.

In one minute. A company called TypeSafe AI has released Jev, an AI model that never writes a sentence. Ask it a question with a fixed set of answers and it returns one answer plus a probability for each option. For a model of this capability class that is new, and it sharpens a question that classifiers never forced on us: how do you check a witness who cannot explain itself? TypeSafe’s answer is to grade Jev against the average opinion of two other AI models. We opened the twenty cases TypeSafe publishes to illustrate its benchmark and found that, in eight of the nineteen where both graders answered, those two models disagreed with each other about the right answer. This note explains why that matters far beyond one company.

1. A model that does not talk

For three years, when a frontier AI model gave us an answer, it gave us words. We could read them, argue with them, and catch the model contradicting itself. Smaller classifiers have always returned bare labels; what is new is a frontier-capability model that returns nothing else. TypeSafe AI, a San Francisco company with no relation to the TypeScript programming language, has shipped a model that does none of that. It is called Jev, and TypeSafe calls it the first of a new category, “System One models”, after the fast, intuitive kind of thinking in Daniel Kahneman’s Thinking, Fast and Slow.

Here is how Jev works. You give it some material, say a customer’s support ticket. You give it a set of typed questions: Is this customer asking for a refund? (yes or no, with a probability). Which team should handle it? (pick one from a list). How frustrated is the customer, from 0 to 4? (a score). Jev returns the answers, each with a probability distribution, in a fraction of a second. No prose, no explanation, no code. Your software acts on the answers directly.

For automating routine decisions this is attractive, and we are not here to say otherwise. Cutting out the free text removes a whole class of failures, and a probability you can set a threshold on is more useful to a program than a paragraph. What interests us is a side effect. A model that only picks from your list can never tell you the list is missing something. A human reviewer, or a model that writes, can say “you asked the wrong question.” Jev cannot. It can hint that none of the options fit well, because its probabilities will be spread out instead of concentrated, and TypeSafe’s own documentation says as much. But it can never name the option you forgot. Whatever you left out of the question set stays out, for every witness that answers it.

We come at this from an unusual angle. Our working paper treats AI outputs the way scholars treat old manuscripts: as testimony from witnesses who may share sources, copy each other, and make the same mistakes for the same reasons. In that frame, Jev is a witness who has stopped writing.

2. Who grades the grader?

Every AI model gets a benchmark score. The question worth asking of any score is: compared to what? For a maths problem the answer is easy: the correct number. For “should this security alert be escalated?” there is no answer key. Somebody has to decide what the right answer was.

TypeSafe publishes its benchmark at evals.typesafe.ai. Its methodology section explains how it decided:

“Instead of debating the correctness of the harness and labels, we assume that the code is correct, and measure against the current smartest large models. For this eval, the reference labels are generated via an average of the responses of GPT-6 Astra and Claude Fable 5.1, both at high thinking, answering every question in the harness. All other models are evaluated using the provider’s default reasoning settings.”

In plain terms: the “right answer” to each question is whatever two leading AI models, OpenAI’s GPT-6 Astra and Anthropic’s Claude Fable 5.1, said on average. Jev’s accuracy is how often it agrees with that average. So is everyone else’s.

TypeSafe is upfront about the obvious problem. The launch post says this “biases answers towards OpenAI and Anthropic’s models. We likely underestimate the relative performance of our model and DeepSeek’s models.” That is an honest caveat. But there is a deeper problem it does not name.

When two AI models are both wrong, they tend to be wrong in the same way. A 2025 study of more than 350 models found that when two models both miss a question, they give the same wrong answer far more often than chance would predict. The effect gets stronger as models get more capable, and it holds across different companies and architectures (Kim, Garg, Peng and Garg, ICML 2025). Averaging two judges dampens the mistakes they make separately; at the level of a final decision it simply picks one judge’s side, as we show below. It does nothing about the mistakes they make together, and at the top of the field, those shared mistakes are the big ones.

So an average of two frontier models is not a neutral referee. It is their shared opinion, promoted to ground truth. A model that copies their shared mistake scores as correct. A model that gets it right when they are both wrong scores as a failure. We called this the consensus trap when we wrote about benchmarks in August. TypeSafe’s benchmark is the clearest disclosed example we have seen. One honesty note of our own: the study measured this on two public leaderboards and a hiring task. That the same thing happens inside a workflow benchmark with no independent truth is our inference, not the study’s finding.

3. We opened the cases. Here is what we found.

The benchmark dashboard scores nine models on four workflows: security alerts, support-agent monitoring, invoice processing, and customer service, 711 cases in all. Jev lands in the middle of the pack: 67.8% agreement with the reference on average, against 74.1% for the best model. It sits below the leading OpenAI and Anthropic models, a tenth of a point behind GPT-5.6 Terra, level with Claude Sonnet 5, and above both DeepSeek models. If you came here asking whether Jev “hallucinates”, TypeSafe’s launch post is candid on that too: the 0% it plots is, in the post’s words, “not empirical. Schema matching is guaranteed, thus we can confidently add 0% into the plots.” Jev cannot return a malformed answer. That is a different thing from returning a correct one.

The part we care about is smaller and more revealing. TypeSafe also publishes twenty worked cases, five per workflow, chosen to illustrate five situations: each of three compared models disagreeing with the other two in turn, all three missing the reference, and all three agreeing. For nineteen of those cases it publishes what each of the two grading models said, question by question, not just their average; in one of the nineteen, most of the second grader’s follow-up answers are missing from the file, and in two of the four workflows the average itself is not published at all. That is more transparency than most benchmarks offer, and it let us count.

The two graders disagreed with each other on the final decision in 8 of the 19 cases. We compared the full list of actions each grader arrived at; in a few cases one grader’s list was the other’s plus extra steps, and we count that as a disagreement too. Question by question, they gave different answers 31 times out of the 356 questions both graders answered, all of which are in the archived files; in six more instances, one of them gave no answer at all. These twenty cases were hand-picked by TypeSafe to show disagreement among the graded models, so this is not a measurement of how often the graders disagree across all 711. It is what TypeSafe chose to publish, read from the other side.

Every single case TypeSafe labels “all three miss the reference” is a case where the two graders also disagreed with each other. Take the security example. One grader said ISOLATE HOST. The other said REVOKE SESSIONS. The average came out as ISOLATE HOST. All three models being graded said QUARANTINE FILE. So three models were marked wrong against an answer that one of the two models writing the answer key did not give either. To be exact about what that shows: the three models would have been marked wrong under either grader’s answer, so this is a contested answer key, not a correct answer being punished. Nobody can say which grader was right: these workflows have no independent answer key, which is the whole point. Two of the four cases labelled “all three agree” cut the other way. In one, the two graders split between ISOLATE HOST and ESCALATE TIER2, the average picked ESCALATE TIER2, and all three graded models said the same, so agreeing with one grader against the other counted as correct. Wherever TypeSafe publishes the average, it equals one grader’s answer or the other’s, never a third. That is not a coincidence: the averaged probabilities are run through the same decision rules as any model’s answers, so the average has to land on some action, and in every published case it lands on one grader’s side.

One of the graders was partly a stand-in. In the security workflow, the file describing the second grader reads, word for word: “Claude Fable 5.1 high, one question per request; Claude Opus 5 high on the 88 documents Fable refused or failed”. So for the documents Fable would not or could not answer, a different Anthropic model, Claude Opus 5, filled in. The file counts 240 cases in that workflow but does not say how documents map onto cases, so we will not turn the 88 into a percentage. What follows is one thing: one grader for one workflow is a blend whose recipe is not disclosed beyond a document count. The file records the stand-in; the leaderboard does not mention it. And because the file does not say which documents Fable failed on, we cannot tell whether that REVOKE SESSIONS above came from Fable or from Opus 5.

Jev’s own answers are published too, and they lean neither way. On the 31 questions where the two graders disagreed, Jev sided with Astra 14 times, with the second grader 14 times, and with neither 3 times. For a 0-to-4 score, TypeSafe’s own convention is the probability-weighted average rather than the single most likely level; scored that way, one answer flips and it becomes 13 to 15. Either way, on this tiny, curated sample there is no sign that Jev leans towards either of the models whose average grades it. But be careful what that shows. An even split is what “no favourite” looks like. It is also, as one of our outside reviewers pointed out, what a model trained to follow the average of the two would look like. This count cannot tell those two stories apart. Hold that thought for the next section.

None of this makes Jev better or worse than the dashboard says. It shows that the benchmark’s definition of “correct” is contested inside its own data, in a way TypeSafe made visible and its headline number does not.

4. Where did Jev learn to judge?

Our research programme asks a question of every model: where did its habits come from? Was it trained on another model’s output? Does it share a base with a competitor? For models that write, there is a surprisingly strong clue in the writing itself: distributions of word choices are distinctive enough to tell five major AI systems apart with 97% accuracy (Sun et al., ICML 2025). It is the machine equivalent of recognising a scribe by their hand.

Jev writes nothing, so that clue is gone. What remains are two harder methods: deliberately planted markers, and comparing a model against an earlier version of itself to see which teacher it learned from. Neither is available to an outsider. TypeSafe’s own primer says its models start from ordinary pretrained language models and add a new training stage, “Reinforcement learning for calibrated decisions”. TypeSafe does address where its training data comes from, in a launch-post FAQ answer that no text render of the page shows and that one of our reviewers found in the page’s script payload: “We make all the data ourselves.” That rules out training on customer data. It does not name a base model, does not name a teacher, and does not say whether self-made data can carry other models’ answers. The same FAQ says Jev is “neither small nor an LLM”, which sits uneasily beside the primer’s description of a third way to adapt pretrained language models; TypeSafe does not reconcile the two, and we cannot. By our own rules, then, Jev’s ancestry is a research programme, not a result, and a thinner one than for models that write.

That leaves one hypothesis we want to state carefully. If Jev was trained to match the same kind of two-model average that its benchmark uses, then a review panel that added Jev as a “fourth independent opinion” would in fact be adding a descendant of two opinions it already had, and any process that counted it as independent would be fooled without knowing. Two things bear on this, neither decisively. Jev is mid-pack on the dashboard, where a model trained to copy the graders might be expected to agree with them more; but ability and agreement are tangled together there, so that is weak evidence. And on the 31 contested questions, Jev split fourteen to fourteen between the two graders (thirteen to fifteen under TypeSafe’s own scoring of one answer type). That is consistent with having no favourite, and equally consistent with tracking the pair’s average, so it settles nothing. We hold the hypothesis as possible and say what would settle it: TypeSafe disclosing where its training answers came from, or a comparison that sets Jev’s closeness to the two graders’ average against other capable models that were not trained on them, measured across all 711 cases rather than the 20 we can see. Closeness alone would not settle it: any capable model tends to land between two capable graders.

5. Our own hand in this

This note was drafted with Claude Fable 5.1, one of the two models whose average TypeSafe treats as truth. An earlier analysis we consulted came from an OpenAI model, the other half of that average. Both authors are, in other words, parties to the thing they were assessing. So before this note was written, we had the underlying assessment attacked by a reviewer from a third company, Moonshot’s Kimi. It corrected three of our quotations and broke nineteen of our conclusions, including a first draft’s claim that Jev “agrees least” with the consensus. It does not; it is mid-pack. A second Kimi seat then recomputed every count above from the raw files and matched them, and broke this note’s first draft in turn: it had converted TypeSafe’s “88 documents” into a share of 240 cases, two units the files do not equate, and had written that a grader was “grading itself”, which the data does not support. Both are gone. A third seat, from xAI’s Grok, recomputed the counts a third time and showed that our security example proved less than we had implied, that placing the stand-in fact next to Opus 5’s ranking invited a conclusion we had disclaimed, and that one of our counts depends on a scoring choice. A fourth, from DeepSeek, recomputed them again and caught us treating an even split as evidence against a hypothesis it cannot distinguish from. A fifth, another Kimi seat, found a FAQ answer about training data that four page-readers, ours included, had reported as absent. A sixth, Alibaba’s Qwen, run by the author rather than by us, confirmed every count and became the fourth to say the stand-in paragraph still insinuated; it now reads as a plain disclosure. All are corrected above. A note about correlated judges written by a correlated judge is worth exactly as much as its outside check.

On 17 September two further reviewers, an OpenAI Codex seat and a Moonshot Kimi seat, examined our plan to test Jev directly and showed that the test we proposed in §4 could not settle the question on its own. The sentence in §4 now says what it would need.

What we are not saying

We are not saying typed decisions are a bad idea; for routing and triage they are probably a good one. We are not saying TypeSafe’s benchmark is dishonest; it discloses more than most, and everything above is counted from what it discloses. We are not saying Jev was trained on its graders’ answers; the one test we could run cannot tell. We are saying three things. A benchmark whose truth is an average of two witnesses inherits whatever those witnesses get wrong together. TypeSafe’s own published cases show the two witnesses disagreeing on the answer in eight of nineteen. And a class of model that cannot write has taken away the evidence we would otherwise use to ask where its judgment came from.

Why we care

CounterProof reviews AI-written code the way an adversary would. Our operating rule is that a finding is settled by running the code, never by taking a vote among models. The new class of decision models is going to be graded, and trained, against model consensus, because consensus is cheap and real outcome labels are expensive. Our whole method is the expensive alternative: frozen cases, reproductions you can execute, and labels that come from what the code actually did rather than from what two models agreed it would do. If typed-decision models become the layer that AI agents act through, then the question of whose judgment that layer carries becomes the provenance question of the decade. The only honest answer will come from checking against outcomes. That is the work.

The technical record. Method, schema, the full counts table, what the specimen does and does not license, and limits: Field Note 1, Consensus as oracle: TypeSafe’s Jev benchmark as a type specimen, on our research site.


Disclosure, drawing nothing from it. Claude Opus 5, the stand-in for part of the security workflow’s answer key, is also one of the nine models being graded. The file does not say which documents were substituted, and a stand-in answer still has to win the average against Astra before it becomes the key. We draw no conclusion about any model’s score from this, and four outside reviewers told us that placing the fact inside the argument above invited one anyway; so it sits here.

Sources. TypeSafe launch post, documentation (introduction, introduction/machine-learning-primer, confidence, System One concepts, HTTP API, primitives/choice) and evals.typesafe.ai, read 16 September 2026; the five pages and four per-case files (*-cases.js) served by that site, archived with SHA-256 hashes in our reports store. Kim, Garg, Peng and Garg, Correlated Errors in Large Language Models, ICML 2025, arXiv:2506.07962. Sun, Yin, Xu, Kolter and Liu, Idiosyncrasies in Large Language Models, ICML 2025, arXiv:2502.12150. Rawat et al., Reference-Based Distillation Detection in LLMs, arXiv:2607.09692. Soons, Agentic Stemmatics, v1.29, doi:10.5281/zenodo.22790100. Counts in §3 are over the twenty published cases only; final decisions were compared as unordered sets of actions; one case carries a single grader’s answer and is excluded from the eight-of-nineteen.

← All notes