Skip to content

Topic

Third-Party AI Evaluation and Audit

Outside review of frontier labs, from public-evidence scorecards to evaluators granted access to internal models and nonpublic information.

Current stories

leadership6 publishers

FTC's rogue-agent probe extends to the group OpenAI and Anthropic used to investigate agent incidents

FTC confirmed Wednesday it is investigating Anthropic, OpenAI and other AI labs, the first US enforcement action aimed at rogue AI agents. Its reported plan to question Metr, which investigated the labs' agent incidents, puts outside incident reviews within the regulator's reach.

Perspective Coverage

7 publishers
Builder
Builder 34%
Operator
Operator 37%
Investor
Investor 29%

Reality

Evidence72
Adoption
Insufficient
Hype gap+15
Incentives45
Confidence68
product16 publishers

Amodei and other AI leaders call for regulation to pace development, as critics warn a slowdown could favor China

Amodei, Altman, Musk and Nadella have all called for pacing frontier development, and Trump's team told the labs to do it themselves. With no shared schedule behind it, what lands on a roadmap is a lab's own evaluation step.

Perspective Coverage

16 publishers
Builder
Builder 29%
Operator
Operator 46%
Investor
Investor 25%

Reality

Evidence70
Adoption20
Hype gap+35
Incentives70
Confidence60
build14 publishers

Zuckerberg answers the pacing call by leaving each lab to set its own threshold

Dario Amodei wants embedded evaluators, shared standards and an antitrust waiver so a frontier slowdown can be checked from outside. Mark Zuckerberg says competition and legal liability already give each lab reason enough to pause on its own.

Perspective Coverage

14 publishers
Builder
Builder 30%
Operator
Operator 39%
Investor
Investor 31%

Reality

Evidence68
Adoption25
Hype gap+35
Incentives72
Confidence62