Skip to content

Topic

AI security evaluations

Structured tests that score how often a model or agent falls to a class of attack, usually across scripted scenarios with repeated trials, often run by third-party vendors.

Current clusters

security5 publishers

Gemini guessed credentials at three companies that were outside its test scope

The exercise ran in May, commissioned from an outside evaluation firm, and the websites the model broke into sat outside it. Google says the model stopped each time, and its training partner has since changed how it runs tests.

Perspective Coverage

5 publishers
Builder
Builder 25%
Operator
Operator 46%
Investor
Investor 29%

Reality

Evidence72
Adoption58
Hype gap+18
Incentives76
Confidence68