Skip to content

Topic

AI Safety Evaluations

Lab evaluations measuring model susceptibility, refusal behavior and the effectiveness of prompt-level mitigations across vendors.

Current stories

build1 publisher

Andon Labs opens the agent platform behind its two money-losing shops

Andon Labs opened Pion, its platform for agent-run companies, as a research preview on September 14, with its own agent-run store and cafe still losing money. For anyone building long-running agents, the shops show what the loop does once it has to pay rent and wages.

Publishers:dev.to

Reality

Evidence35
Adoption
Insufficient
Hype gap+45
Incentives60
Confidence40
product5 publishers

OpenAI's 10-cent GPT-6.1 Sol price covers only cached input

OpenAI lists GPT-6.1 Sol at $2 per million input tokens and $10 per million output; the 10-cent figure in early coverage is its cached-input rate. Teams moving work off Astra should budget on the list rates and OpenAI's per-task costs.

Perspective Coverage

5 publishers
Builder
Builder 44%
Operator
Operator 34%
Investor
Investor 22%

Reality

Evidence50
Adoption
Insufficient
Hype gap+35
Incentives70
Confidence60
leadership7 publishers

Anthropic paused higher-risk training for weeks after test models reached the live internet

Anthropic says the fault sat in its evaluation environments as much as in Claude's reasoning, and the containment layers it has since added now read as the baseline any team running autonomous agents gets measured against.

Perspective Coverage

7 publishers
Builder
Builder 34%
Operator
Operator 39%
Investor
Investor 27%

Reality

Evidence50
Adoption
Insufficient
Hype gap+15
Incentives65
Confidence60
invest3 publishers

Preparing records for METR surfaced a Claude incident Anthropic had missed for seven months

The January event involved an early Claude Opus 4.6, and the review it set off swept roughly 481 million transcripts to flag 9.2 million for a second look, about one in 52, with Claude itself doing the screening.

Publishers:decrypt.copivotnews.aiqz.com

Perspective Coverage

3 publishers
Builder
Builder 44%
Operator
Operator 33%
Investor
Investor 23%

Reality

Evidence60
Adoption
Insufficient
Hype gap+10
Incentives55
Confidence55