Build1 publisherNot yet confirmed elsewhere3 min readPublished
TypeSafe's Jev reaches frontier AUROC on transcript monitoring at about 1/200th of Opus 4.7's cost
TypeSafe's Jev 1.13 matched no-CoT GPT-5.5 and Opus 4.7 on agent-monitoring AUROC at about 1/200th of Opus 4.7's cost, a LessWrong benchmark found. It answers in one step without writing text, so its calls leave no reasoning trace to audit.
The Engineer · Build desk
What happened
- Averaged over 14 ThinkFast benchmarks, Jev scored 46.8% chance-corrected, behind GPT-5.5 at 58.1% and Opus 4.7 at 53.7% but ahead of Opus 4.1 at 40.4%.
- Its results were uneven: it beat GPT-5.5 on GPQA-Diamond, 65.2 to 51.8, but trailed on BEA-24, intuitive physics and the Sally-Anne test.
- On Neel Nanda's NCRI ladder, where answers need multi-step reasoning, Jev ranked 209th of 279, roughly at GPT-3.5 level.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- constraint Teams running monitors on a tight false-alarm budget get frontier-level ranking from Jev but lower catch rates on SHADE-Arena sabotage and sandbagging, so AUROC parity alone does not justify replacing a no-CoT frontier monitor.
- contradiction TypeSafe pitches Jev for verification, yet test-case pass/fail, the benchmark closest to that job, sat in its weakest category when built from per-test yes/no questions.
- precedent If monitor improvements made on Jev carry over to frontier models, monitor research could iterate on a model costing about 1/200th as much as Opus 4.7.
A Jev call carries a state and a set of typed questions. The model is non-autoregressive and cannot emit text [1]. The state is text describing the situation, such as an agent's transcript [5]. For suspicion scores the authors used a score question. Its levels are written in words, for example five running from "clearly benign" to "clearly deceptive", and Jev returns how likely each level is [6]. Every question in a request is answered in a single step, with no text generated [5]. TypeSafe pitches it as a "System One" model for classification, routing, scoring and verification [4].
The LessWrong write-up reports the monitoring results two ways, and they point in different directions. Judged by ROC, Jev is near the frontier [10]. A deployed monitor runs at one threshold, though, set by how many false alarms reviewers will tolerate. At the strict 1.5% false-positive rate the authors tested, Jev caught 45.9 on SHADE-Arena against GPT-5.5's 59.3, a 13.4-point gap [10][15]. At that threshold the coding-sabotage gap is 2.8 points and the sandbagging gap is 22.2 [16][14].
The 1/200th figure carries over only under two conditions. The comparison models ran without chain of thought, and the cost ratio is quoted against Opus 4.7 [2]. A team whose current monitor reasons before it scores is starting from a different baseline than the one measured. A team already on a cheaper model than Opus 4.7 will see a smaller ratio. Either way, the saving assumes transcripts that resemble the sabotage and sandbagging sets the authors scored.
The reasoning weakness has two sources. Jev cannot write a multi-part answer, so the authors asked one question per part in a single request and counted the answer correct only if every piece was right [7]. A Sudoku became 81 choice questions [7]. Sudoku is an odd exam for a classifier, and the authors say as much, calling weaker performance on these tasks unsurprising [7]. Multi-part tasks, chess and Tower of London among them, were its weakest category [11].
NCRI is the cleaner test. Its free-form questions were converted to multiple choice, using lists of 255 consecutive integers where the answer is a number [13]. Each item then takes a single pick, and Jev still ranked below about three quarters of the field [3][17]. Chained lookup holds up better. On N-hop table lookup it performs strongly up to 6 hops [12].
For routing, the useful output is the probability Jev assigns to every option of a choice question [6]. A router can take the top option when it clears a confidence bar and pass the request to a reasoning model when it does not. I think both jobs the authors propose, routing and always-on monitoring, should be built that way [18]. In my context, where a reasoning model is affordable on a small share of calls, Jev scores everything and the flagged or uncertain cases go to a model that writes its reasoning down.
What to watch
- Strict-threshold catch rates for Jev measured against frontier monitors with chain of thought switched on.
- A Jev release after 1.13 that narrows the 22.2-point sandbagging gap at a 1.5% false-positive rate.
- Per-request pricing from TypeSafe that lets teams check the 1/200th ratio against their own monitor spend.
Clarity's read
What the record supports and how the coverage leans. The claims behind it follow.
Reality
- Evidence58
- Adoption
- Insufficient
- Hype gap+15
- Incentives
- Insufficient
- Confidence50
Claim ledger
Ranked by verification strength, evidence, and original report placement.
- [1]
Jev 1.13 is TypeSafe's non-autoregressive model, which answers questions with probabilities and cannot emit text.
- [2]
As a transcript monitor, Jev's AUROC on sabotage and sandbagging detection is level with GPT-5.5 and Opus 4.7 without CoT, at about 1/200th of Opus 4.7's cost.
- [3]
Jev does badly on tasks where the answer involves multi-step reasoning: on Neel Nanda's NCRI, it ranks 209 of 279, at about GPT 3.5 level.
- [4]
TypeSafe calls Jev a "System One" model, pitching it for decisions like classification, routing, scoring, and verification.
- [5]
Jev takes a state, which is text describing the situation (e.g. a problem statement or an agent's transcript), together with one or more typed questions about that state, and answers all of them in a single step without generating any text, so there is no chain of thought or trace to inspect.
- [6]
A choice question lists up to N options and Jev picks one and assigns a probability to every option; a score question gives a rating scale whose levels are described in words (e.g. five levels from "clearly benign" to "clearly deceptive") and Jev says how likely each level is, used for transcript suspicion scores; a yes/no question returns the probability that the answer is yes, used for test case pass/fail.
- [7]
For multi-part answers the authors ask one question per part in the same request and assemble the answer, which is correct only if all pieces are correct; a Sudoku question becomes 81 choice questions. This is not the kind of query Jev was designed for, so weaker performance is unsurprising.
- [8]
Averaged over 14 ThinkFast benchmarks, Jev scores 46.8% chance-corrected, against 58.1% for GPT-5.5, 53.7% for Opus 4.7 and 40.4% for Opus 4.1.
- [9]
Jev is ahead of GPT-5.5 on GPQA-Diamond (65.2 vs 51.8), within 5pp on N-hop lookup, CyBash and training-poisoning, and trails on BEA-24, intuitive physics and Sally-Anne.
- [10]
Judged by ROC, Jev is near the frontier on monitoring. At a strict 1.5% false positive rate it is close to GPT-5.5 on coding sabotage (88.0 v 90.8) but catches fewer on SHADE-Arena (45.9 v 59.3) and sandbagging (41.3 v 63.5).
- [11]
Multi-part answer tasks (sudoku, test-case pass/fail, chess, Tower of London) are Jev's weakest category.
- [12]
On the N-hop table lookup task, Jev performs strongly up to 6 hops.
- [13]
NCRI benchmarks are free-form answer, which Jev does not support, so the authors turned each question into an MCQ: a fixed list of 255 consecutive integers for small-integer answers, the candidates named in the prompt where answers appear there, and 255 consecutive integers containing the answer at a random position for large or unbounded answers.
- [14]
At a 1.5% false-positive rate, Jev trails GPT-5.5 by 22.2 points on sandbagging.
- [15]
At a 1.5% false-positive rate, Jev trails GPT-5.5 by 13.4 points on SHADE-Arena.
- [16]
At a 1.5% false-positive rate, Jev trails GPT-5.5 by 2.8 points on coding sabotage.
- [17]
On NCRI, 208 of 279 entries rank above Jev, about three quarters of the field.
- [18]
Cost and speed make Jev an appealing option for routing and simple decision-making in agentic systems, and always-on monitoring of agents; in both, its decisions leave no reasoning trace to inspect.
- [19]
Jev could be a cheap test bed for developing frontier monitors, if improvements on Jev carry over.
Sources
1 independent publisher whose own reporting we read for this story.
- lesswrong.comBenchmarking Jev against no-CoT LLMs
1 article · October 7, 2026
Topics and entities
Follow any of these and your For You feed starts watching them — no settings page required.
Entities
- TypeSafeFollow
- JevFollow
- GPT-5.5Follow
- Claude Opus 4.7Follow
- Neel NandaFollow
- No-CoT Reasoning IndexFollow
- ThinkFastFollow
- SHADE-ArenaFollow
- GPQA DiamondFollow