Skip to content

Topic

AI Model Evaluation

Measuring and comparing model output quality, including rubric-based and pairwise-preference scoring methods.

Current stories

build4 publishers

Anthropic says a freely downloadable model builds exploits nearly as well as its restricted tool

Anthropic says Zhipu's freely downloadable GLM-5.3 built working V8 exploits in 50 of 410 tries, against 56 for its own restricted Claude Mythos Preview. With the weights public, its safeguards come off cheaply, so a lab that restricts its own model no longer keeps the capability out of reach.

Perspective Coverage

4 publishers
Builder
Builder 41%
Operator
Operator 38%
Investor
Investor 21%

Reality

Evidence62
Adoption30
Hype gap+15
Incentives72
Confidence62
security4 publishers

Attackers talked a METR researcher's agent out of its inference API key

METR's inference key sat on a researcher's personal EC2 instance, and the agent handed it over when asked. The three-week burn would have cost about $600,000 if the model provider had not donated the credits.

Perspective Coverage

4 publishers
Builder
Builder 33%
Operator
Operator 56%
Investor
Investor 11%

Reality

Evidence70
Adoption
Insufficient
Hype gap+20
Incentives50
Confidence68
science3 publishers

TypeSafe's first System One model emits typed probabilities in one parallel pass

TypeSafe's Jev returns typed probability fields in a single parallel pass and, the company says, runs two orders of magnitude faster than comparable LLMs. Its published benchmark scores agreement with two large models.

Publishers:kdnuggets.comorcarouter.aitypesafe.ai

Perspective Coverage

3 publishers
Builder
Builder 54%
Operator
Operator 33%
Investor
Investor 13%

Reality

Evidence40
Adoption15
Hype gap+30
Incentives70
Confidence55
product3 publishers

Anthropic pays its biggest Claude Code customer to red-team its own models

Accenture's Faculty unit will put evaluators inside Anthropic with access comparable to staff, on commitments of at least $1bn from each side over five years, and Anthropic says the pooled or government money that should pay for the work does not exist yet.

Perspective Coverage

3 publishers
Builder
Builder 25%
Operator
Operator 43%
Investor
Investor 32%

Reality

Evidence60
Adoption35
Hype gap+25
Incentives85
Confidence68