Build2 distinct publishers3 min readUpdated
Guidelight's first control assessment puts Anthropic and OpenAI at C+, Google at D+, xAI at D-, and Meta at F, using public evidence only. That is a baseline, not a lab's own account.
The Engineer · Build desk
Compiled by The EngineerSomething wrong?How this is made
Guidelight AI Standards, a new independent nonprofit, has published the first outside grading of how frontier labs control the AI systems running inside their own companies [1]. The top grade was a C+ [2], which matters less as a verdict on any one lab than as a number a buyer can put next to a vendor's description of its own internal agent oversight.
The group scored Anthropic, OpenAI, Google, xAI and Meta on six practices drawn from its broader Control standard: logging internal AI activity, measuring whether monitors actually catch misbehavior, gating high-risk actions before they execute, halting systems after a burst of suspicious activity, commissioning outside reviews, and keeping a containment plan for a model that tries to defeat its controls [5]. Anthropic and OpenAI tied at C+, Google took a D+, xAI a D-, and Meta an F [2]. Scores ran from 0 for not implemented to 5 for full implementation, with 2 meaning limited partial implementation and 3 meaning substantial partial implementation; each practice counted equally [6].
The aggregate is thin. Across the five companies the average rating was 1.6 out of 5, 22 of the 30 individual ratings (73 percent) were a 2 or lower, seven were zeros, and no company earned a 4 or 5 on any practice; those totals are theneuron.ai's calculations from Guidelight's published scorecard [7]. An average of 1.6 is 32 percent of the available points [16]. Anthropic scored a 3 on five of six practices; OpenAI recorded three 3s and three 2s [8]. Since no rating exceeded 3 and 22 ratings were 2 or lower, there were eight 3s in total, and Anthropic's five plus OpenAI's three account for all of them, meaning Google, xAI and Meta did not clear a 2 anywhere [15].
Where progress exists, it is concentrated in detection rather than prevention. Anthropic and OpenAI both took 3s for logging and monitor efficacy, with Guidelight finding they log substantial portions of internal agent activity, scan it for suspicious behavior, and have done some work measuring how often the monitors catch what they should [9]. The pattern across all five is best at spotting misbehavior, worst at prevention and containment [10]. Google has published a detailed control roadmap [11].
Independent work points the same way. In a frontier-risk pilot earlier this year, the research group METR was given access to internal models and nonpublic information at Anthropic, Google, Meta and OpenAI, and found that automated monitors caught many harmful actions but also had coverage gaps and could be disabled or evaded, particularly when Anthropic let an embedded researcher red-team part of the system [12]. xAI is the one graded company not among METR's four participants [18].
The caveats are real and Guidelight states them. It is not a regulator, this is not an audit, the grading used only public material such as system cards, safety reports and blog posts, and the standards are living documents the group expects to revise [13][4]. It was founded by two former OpenAI safety leads, Page Hedley and Steven Adler, and Adler discloses that he still holds OpenAI equity [3][13].
For procurement, the six practices are the usable output: ask what fraction of agent activity is logged, what the measured monitor catch rate is, which actions are gated pre-execution, what trips a shutdown, who has reviewed it from outside, and what the containment plan says. Because the grades rest on public evidence [4], a vendor that thinks its score is unfair has an obvious remedy, which is exactly the disclosure a buyer wants.
Follow any of these and your For You feed starts watching them — no settings page required.
Ranked by verification strength, evidence, and original report placement.
Guidelight AI Standards graded Anthropic, OpenAI, Google, xAI and Meta across six safeguards for AI systems used inside their own companies, in its first such assessment.
Anthropic and OpenAI tied at the top with C+ grades, followed by Google at D+, xAI at D-, and Meta at F.
Guidelight is a new independent nonprofit founded by Page Hedley and Steven Adler, two former OpenAI safety leaders. Hedley, the CEO, was previously OpenAI's policy and ethics adviser; Adler, chief scientist, worked at OpenAI from 2020 to 2024 and led teams including product safety and dangerous-capability evaluations.
Guidelight drew only on public sources such as system cards, safety reports and blog posts.
Guidelight selected six practices from its broader Control standard: logging internal AI activity, measuring whether monitors catch misbehavior, gating high-risk actions before they happen, halting systems after a burst of suspicious activity, commissioning outside reviews, and maintaining a containment plan for a model that tries to defeat its controls.
Across all five companies the average rating was 1.6 out of 5; 22 of the 30 individual ratings, or 73%, were a 2 or lower; seven received a zero; and no company earned a 4 or 5 on any practice. The publication states these are its own calculations from Guidelight's scorecard.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Specific and internally consistent, but one detailed source and no primary scorecard
The cluster supplies granular, checkable figures - a 0-to-5 rubric, per-practice scores, category averages, an aggregate of 1.6 out of 5 - and one publisher openly flags that the aggregate statistics are its own calculations from Guidelight's scorecard. Corroboration from METR's pilot strengthens the underlying monitoring claims. Weaknesses: only theneuron.ai carries the detail, the Guidelight scorecard itself is not among the supplied sources, all grading rests on public documents rather than inspection, and no graded company responds.
Partial control implementation at labs; no evidence yet of scorecard uptake
Adoption of the practices being graded is measurable and low-to-partial: an average of 1.6 out of 5 (about 32% of available points), 22 of 30 ratings at 2 or lower, seven zeros, and no rating above 3 anywhere. The clearest real uptake is third-party access - four of five labs joined METR's pilot - and monitoring/logging at Anthropic and OpenAI. Adoption of Guidelight's standard itself is unevidenced: neither source shows a customer, investor or regulator citing the scorecard.
Slightly overstated: letter grades and 'failing' framing outrun what public-evidence scoring can show
The grades are real and specific, but two framings run ahead of the evidence. First, 'failing to keep their own systems in check' compresses scores that the more detailed source explicitly reads as work underway - partial implementation, not absence - and public-evidence-only scoring penalizes non-disclosure as if it were non-implementation. Second, the premise that buyers 'can now cite' the assessment is not evidenced: no source documents any customer, investor or regulator using it, and Guidelight is expressly neither regulator nor auditor with a disclosed equity conflict. Offsetting this, the underlying numbers are conservative and METR's independent findings point the same direction, so the gap is modest rather than large.
Meaningful disclosed conflicts, partially mitigated
Guidelight is founded and run by two former OpenAI safety leaders, and its chief scientist retains OpenAI equity - directly relevant when OpenAI ties for the top grade. Mitigations are stated: the nonprofit says it takes no money from AI companies or their staff, publishes its scale and practices, and calls the standards revisable. On the publishing side, one source is a short brief that ends in a subscription solicitation, and neither source obtained comment from a graded company, so the labs' own incentives are unrepresented in the cluster.
Solid on what was graded, weaker on what the grades mean
Two independent publishers agree on the grades, the five companies and the detection-over-prevention pattern, and the detailed source discloses its own arithmetic and the grader's conflicts - all of which raise confidence in the reported facts. Confidence is capped by dependence on a single detailed source for the numeric breakdown, absence of the primary scorecard, no lab response, public-evidence-only methodology, and no evidence about downstream use of the assessment.
invest
The labs got better at watching their agents escape. They did not get better at stopping them.1 distinct publisher
build
First-turn evals test the safest part of your product, a 90,000-exchange audit finds1 distinct publisher
build
Grok 4.6 lands in Copilot two days after launch, and the model picker becomes a procurement problem1 distinct publisher
product
Washington's secret AI test is coming for open weights, and release dates go with it2 distinct publishers
Distinct publishers with included, body-backed reporting in this cluster.
1 article · August 19, 2026
1 article · August 19, 2026