Skip to content

Build1 publisherNot yet confirmed elsewhere3 min readPublished

TypeSafe's Jev reaches frontier AUROC on transcript monitoring at about 1/200th of Opus 4.7's cost

TypeSafe's Jev 1.13 matched no-CoT GPT-5.5 and Opus 4.7 on agent-monitoring AUROC at about 1/200th of Opus 4.7's cost, a LessWrong benchmark found. It answers in one step without writing text, so its calls leave no reasoning trace to audit.

The Engineer · Build desk

How we use AISend a correction

Photograph accompanying TypeSafe's Jev reaches frontier AUROC on transcript monitoring at about 1/200th of Opus 4.7's cost
Photo: lesswrong.com

What happened

  • Averaged over 14 ThinkFast benchmarks, Jev scored 46.8% chance-corrected, behind GPT-5.5 at 58.1% and Opus 4.7 at 53.7% but ahead of Opus 4.1 at 40.4%.
  • Its results were uneven: it beat GPT-5.5 on GPQA-Diamond, 65.2 to 51.8, but trailed on BEA-24, intuitive physics and the Sally-Anne test.
  • On Neel Nanda's NCRI ladder, where answers need multi-step reasoning, Jev ranked 209th of 279, roughly at GPT-3.5 level.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • constraint Teams running monitors on a tight false-alarm budget get frontier-level ranking from Jev but lower catch rates on SHADE-Arena sabotage and sandbagging, so AUROC parity alone does not justify replacing a no-CoT frontier monitor.
  • contradiction TypeSafe pitches Jev for verification, yet test-case pass/fail, the benchmark closest to that job, sat in its weakest category when built from per-test yes/no questions.
  • precedent If monitor improvements made on Jev carry over to frontier models, monitor research could iterate on a model costing about 1/200th as much as Opus 4.7.

A Jev call carries a state and a set of typed questions. The model is non-autoregressive and cannot emit text [1]. The state is text describing the situation, such as an agent's transcript [5]. For suspicion scores the authors used a score question. Its levels are written in words, for example five running from "clearly benign" to "clearly deceptive", and Jev returns how likely each level is [6]. Every question in a request is answered in a single step, with no text generated [5]. TypeSafe pitches it as a "System One" model for classification, routing, scoring and verification [4].

The LessWrong write-up reports the monitoring results two ways, and they point in different directions. Judged by ROC, Jev is near the frontier [10]. A deployed monitor runs at one threshold, though, set by how many false alarms reviewers will tolerate. At the strict 1.5% false-positive rate the authors tested, Jev caught 45.9 on SHADE-Arena against GPT-5.5's 59.3, a 13.4-point gap [10][15]. At that threshold the coding-sabotage gap is 2.8 points and the sandbagging gap is 22.2 [16][14].

The 1/200th figure carries over only under two conditions. The comparison models ran without chain of thought, and the cost ratio is quoted against Opus 4.7 [2]. A team whose current monitor reasons before it scores is starting from a different baseline than the one measured. A team already on a cheaper model than Opus 4.7 will see a smaller ratio. Either way, the saving assumes transcripts that resemble the sabotage and sandbagging sets the authors scored.

The reasoning weakness has two sources. Jev cannot write a multi-part answer, so the authors asked one question per part in a single request and counted the answer correct only if every piece was right [7]. A Sudoku became 81 choice questions [7]. Sudoku is an odd exam for a classifier, and the authors say as much, calling weaker performance on these tasks unsurprising [7]. Multi-part tasks, chess and Tower of London among them, were its weakest category [11].

NCRI is the cleaner test. Its free-form questions were converted to multiple choice, using lists of 255 consecutive integers where the answer is a number [13]. Each item then takes a single pick, and Jev still ranked below about three quarters of the field [3][17]. Chained lookup holds up better. On N-hop table lookup it performs strongly up to 6 hops [12].

For routing, the useful output is the probability Jev assigns to every option of a choice question [6]. A router can take the top option when it clears a confidence bar and pass the request to a reasoning model when it does not. I think both jobs the authors propose, routing and always-on monitoring, should be built that way [18]. In my context, where a reasoning model is affordable on a small share of calls, Jev scores everything and the flagged or uncertain cases go to a model that writes its reasoning down.

What to watch

  • Strict-threshold catch rates for Jev measured against frontier monitors with chain of thought switched on.
  • A Jev release after 1.13 that narrows the 22.2-point sandbagging gap at a 1.5% false-positive rate.
  • Per-request pricing from TypeSafe that lets teams check the 1/200th ratio against their own monitor spend.

Clarity's read

What the record supports and how the coverage leans. The claims behind it follow.

Reality

Evidence58
Adoption
Insufficient
Hype gap+15
Incentives
Insufficient
Confidence50
Why these scores

Claim ledger

Ranked by verification strength, evidence, and original report placement.

  1. [1]

    Jev 1.13 is TypeSafe's non-autoregressive model, which answers questions with probabilities and cannot emit text.

    ReportedSupportedSource: LessWrong benchmark authorsView cited source
  2. [2]

    As a transcript monitor, Jev's AUROC on sabotage and sandbagging detection is level with GPT-5.5 and Opus 4.7 without CoT, at about 1/200th of Opus 4.7's cost.

    ReportedSupportedSource: LessWrong benchmark authorsView cited source
  3. [3]

    Jev does badly on tasks where the answer involves multi-step reasoning: on Neel Nanda's NCRI, it ranks 209 of 279, at about GPT 3.5 level.

    ReportedSupportedSource: LessWrong benchmark authorsView cited source

Sources

1 independent publisher whose own reporting we read for this story.

  1. lesswrong.com

    1 article · October 7, 2026

    Benchmarking Jev against no-CoT LLMs

Share your take

Let Clarity write the post for you.

Signed-in readers get a short post drafted on this story in the register they choose — narrative, analytical, or a direct position — editable to the last word before it goes anywhere. The share buttons at the top of this story work without an account.

Topics and entities

Follow any of these and your For You feed starts watching them — no settings page required.

Topics

Loading related stories