Security1 publisher2 min readPublished
Sophos wants SOCs to test TypeSafe's Jev for accuracy and calibration before it closes alerts
Sophos cites an independent test putting TypeSafe's Jev at 83% on Banking77 with no training examples, 10 points behind a trained classifier. Jev costs about a twelfth as much per email as Claude Haiku 4.5, so SOCs will be tempted to automate at a volume where small error rates add up.
The Watch · Security desk

What happened
- TypeSafe.ai's Jev takes structured context, a question and a fixed list of permitted answers, and returns probabilities over those answers without writing a long-form text response.
- Natural-language inference models have done zero-shot classification since at least 2019 by testing whether an input supports a statement such as "this alert is malicious".
- On Banking77 the same evaluation measured Jev's median response time at 0.44 seconds, against 1.51 seconds for GPT-5.6 Terra.
Compiled by The WatchSomething wrong?How this is made
Why it matters
- exposure At a twelfth of the price per decision, a SOC can send many more alerts down the close-as-benign path for the same spend, and every wrong close on a real intrusion is one no analyst reviews.
- decision A SOC with a labeled alert history should try a supervised classifier first. Jev makes the most sense where triage questions change often or no labels exist.
- constraint Benchmark scores cannot stand in for triage results. A SOC has to measure on its own alerts how often Jev picks the right action and whether its confidence flags its mistakes.
The close is the action that matters. Sophos frames the job with a PowerShell alert for a script that downloads and runs a file. The agent picks one of three moves: close it as benign, collect more evidence, or send it to an analyst [1]. A wrong escalation costs an analyst some time. A wrong close on a real intrusion means no human sees it [1].
Sophos separates three properties. Always returning a valid answer makes Jev easy to integrate. How often it picks the correct answer decides whether it is useful. Knowing when a proposed answer is unreliable decides how much work can be safely automated [8]. The attention around the launch "has blurred them a bit," Sophos wrote [9].
TypeSafe pitches Jev as much faster and cheaper per decision than a general-purpose LLM [4]. Cofounder and CEO Diogo Almeida said in a Latent Space interview that for this class of models "the goal is for code to be the consumer" [5]. He wants the judgments reliable enough that developers call them as routinely as a database query [20]. The company named Jev after Jevons paradox, in which greater efficiency creates enough new demand to raise total consumption [6]. Sophos applies the paradox to errors. If decisions get cheap enough to run on everything, teams will automate far more of them, and any error rate above the current human or automated baseline adds up to more total mistakes [7].
On the 200 CLINC150 requests, Jev got 26 wrong, GPT-5.6 Terra 16 and gpt-5.4-nano 40 [1]. Terra sits below GPT-5.6 Sol and GPT-6 Astra in OpenAI's lineup [13]. On Banking77, Jev's 17% error rate is about 2.4 times the trained classifier's 7% [2]. At a price gap of about 12 times per email, the same budget buys about 12 times as many decisions. Hold the error rate fixed and the number of wrong ones rises by the same factor [3]. Sophos adds that more model responses can compound mistakes, depending on how the answers are combined [18].
Sophos advises teams that have training and evaluation data to start with a supervised classifier before trusting a zero-shot model with important decisions [15]. Teams whose questions change often, or who have no labeled data, have fewer options [15]. TypeSafe has not published Jev's architecture [10]. According to Sophos, its stated training goal of calibrated decisions resembles published research [19].
The accuracy figures come from a chatbot intent benchmark and a banking support benchmark. The phishing test is cited for its price comparison [12][14][17]. Neither accuracy test used security alerts [12][14].
What to watch
- Whether TypeSafe publishes Jev's architecture or calibration data showing that its probabilities drop on the answers it gets wrong.
- An independent Jev evaluation on labeled SOC alerts that reports close-as-benign errors separately from escalation errors.
- Security vendors wiring Jev or similar fixed-answer models into automatic alert closure.