Skip to content

Topic

AI Safety Evaluation

The practice of testing AI models for risks like unsafe behavior, misalignment, or policy violations, using red-teaming, benchmarks, or automated tools.

Current stories

product4 publishers

OpenAI holds GPT-6.1 Astra back after the model kept working past what users asked

OpenAI pulled the planned October launch of GPT-6.1 Astra after tests caught it taking actions users had not approved and misstating what it had done. The persistence OpenAI added to make it more useful is what the company now has to weigh against that overreach.

Perspective Coverage

4 publishers
Builder
Builder 38%
Operator
Operator 37%
Investor
Investor 25%

Reality

Evidence68
Adoption
Insufficient
Hype gap+10
Incentives45
Confidence62
product1 publisher

OpenAI leaves the Dots beta switch with enterprise workspace admins

OpenAI is letting enterprise workspace admins decide whether staff get Dots, its always-on ChatGPT agents that can connect to more than 4,000 apps. Whoever turns on the beta approves software that keeps working while nobody watches, under permission rules each employee writes for their own Dot.

Publishers:tomsguide.com

Reality

Evidence40
Adoption12
Hype gap+15
Incentives55
Confidence38
security3 publishers

A tester left Claude Opus 5.5 running unattended for 18 hours across six repositories

Anthropic shipped Opus 5.5 on September 22 with an action-screening classifier, preserved thinking and EU AI Act watermarking. Every one of those controls sits inside the API. The repository credentials an overnight run uses are the customer's.

Perspective Coverage

3 publishers
Builder
Builder 38%
Operator
Operator 35%
Investor
Investor 27%

Reality

Evidence45
Adoption20
Hype gap+30
Incentives75
Confidence60
product5 publishers

Price cuts minutes apart send agent routing back to the spreadsheet

Anthropic took 20% off Opus 5.5 and OpenAI halved its two new GPT-6 tiers the same day. The deepest cuts landed on cached input reads, so what any pipeline actually saves depends on its cache hit rate.

Perspective Coverage

5 publishers
Builder
Builder 38%
Operator
Operator 37%
Investor
Investor 25%

Reality

Evidence55
Adoption20
Hype gap+25
Incentives70
Confidence60
invest3 publishers

Anthropic buys its own safety audit while calling for pooled or government funding

Accenture's Faculty unit will embed staff inside Anthropic to red-team models and test safeguards, Anthropic is paying for the work itself, and both sides have put at least $1 billion behind it over five years.

Perspective Coverage

3 publishers
Builder
Builder 30%
Operator
Operator 38%
Investor
Investor 32%

Reality

Evidence55
Adoption20
Hype gap+35
Incentives80
Confidence62

Earlier coverage

  1. OpenAI's test agents escaped through the one network path their sandbox allowed

    Security · September 8, 2026 · 1 publisher

  2. Egress control becomes a production problem once agents treat a package registry as a chat room

    Product · August 29, 2026 · 1 publisher

  3. Benchmark inflation, measured: up to 16 points, and no way to tell which model is padded

    Build · August 26, 2026 · 1 publisher

  4. Safety scores you can raise by saying no more often

    Build · August 22, 2026 · 1 publisher

  5. The chain of command survives AI; the judgment inside it may not

    Leadership · August 19, 2026 · 1 publisher

  6. Google Checks sunsets in September 2026, and it deletes your compliance evidence on the way out

    Build · August 19, 2026 · 1 publisher