Skip to content

Build2 publishers3 min readPublished

The best grade for controlling in-house AI agents is a C+, and buyers can now cite it

Guidelight's first control assessment puts Anthropic and OpenAI at C+, Google at D+, xAI at D-, and Meta at F, using public evidence only. That is a baseline, not a lab's own account.

The Engineer · Build desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened

  • Guidelight AI Standards graded Anthropic, OpenAI, Google, xAI and Meta across six safeguards for AI systems used inside their own companies, in its first such assessment.
  • Anthropic and OpenAI tied at the top with C+ grades, followed by Google at D+, xAI at D-, and Meta at F.
  • Guidelight is a new independent nonprofit founded by Page Hedley and Steven Adler, two former OpenAI safety leaders. Hedley, the CEO, was previously OpenAI's policy and ethics adviser; Adler, chief scientist, worked at OpenAI from 2020 to 2024 and led teams including product safety and dangerous-capability evaluations.
  • Guidelight drew only on public sources such as system cards, safety reports and blog posts.
  • Guidelight selected six practices from its broader Control standard: logging internal AI activity, measuring whether monitors catch misbehavior, gating high-risk actions before they happen, halting systems after a burst of suspicious activity, commissioning outside reviews, and maintaining a containment plan for a model that tries to defeat its controls.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

Guidelight AI Standards, a new independent nonprofit, has published the first outside grading of how frontier labs control the AI systems running inside their own companies [1]. The top grade was a C+ [2], which matters less as a verdict on any one lab than as a number a buyer can put next to a vendor's description of its own internal agent oversight.

The group scored Anthropic, OpenAI, Google, xAI and Meta on six practices drawn from its broader Control standard: logging internal AI activity, measuring whether monitors actually catch misbehavior, gating high-risk actions before they execute, halting systems after a burst of suspicious activity, commissioning outside reviews, and keeping a containment plan for a model that tries to defeat its controls [5]. Anthropic and OpenAI tied at C+, Google took a D+, xAI a D-, and Meta an F [2]. Scores ran from 0 for not implemented to 5 for full implementation, with 2 meaning limited partial implementation and 3 meaning substantial partial implementation; each practice counted equally [6].

The aggregate is thin. Across the five companies the average rating was 1.6 out of 5, 22 of the 30 individual ratings (73 percent) were a 2 or lower, seven were zeros, and no company earned a 4 or 5 on any practice; those totals are theneuron.ai's calculations from Guidelight's published scorecard [7]. An average of 1.6 is 32 percent of the available points [16]. Anthropic scored a 3 on five of six practices; OpenAI recorded three 3s and three 2s [8]. Since no rating exceeded 3 and 22 ratings were 2 or lower, there were eight 3s in total, and Anthropic's five plus OpenAI's three account for all of them, meaning Google, xAI and Meta did not clear a 2 anywhere [15].

Where progress exists, it is concentrated in detection rather than prevention. Anthropic and OpenAI both took 3s for logging and monitor efficacy, with Guidelight finding they log substantial portions of internal agent activity, scan it for suspicious behavior, and have done some work measuring how often the monitors catch what they should [9]. The pattern across all five is best at spotting misbehavior, worst at prevention and containment [10]. Google has published a detailed control roadmap [11].

Independent work points the same way. In a frontier-risk pilot earlier this year, the research group METR was given access to internal models and nonpublic information at Anthropic, Google, Meta and OpenAI, and found that automated monitors caught many harmful actions but also had coverage gaps and could be disabled or evaded, particularly when Anthropic let an embedded researcher red-team part of the system [12]. xAI is the one graded company not among METR's four participants [18].

The caveats are real and Guidelight states them. It is not a regulator, this is not an audit, the grading used only public material such as system cards, safety reports and blog posts, and the standards are living documents the group expects to revise [13][4]. It was founded by two former OpenAI safety leads, Page Hedley and Steven Adler, and Adler discloses that he still holds OpenAI equity [3][13].

For procurement, the six practices are the usable output: ask what fraction of agent activity is logged, what the measured monitor catch rate is, which actions are gated pre-execution, what trips a shutdown, who has reviewed it from outside, and what the containment plan says. Because the grades rest on public evidence [4], a vendor that thinks its score is unfair has an obvious remedy, which is exactly the disclosure a buyer wants.

Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories