Science1 distinct publisher3 min readPublished
Internal telemetry reported by Reuters shows code changes up 220% while user-facing changes rose 36%, and major incidents up 40%. The cuts of up to 60% did not happen.
The Scientist · Science desk

Compiled by The ScientistSomething wrong?How this is made
Put Bosworth's two figures over each other. Internal code changes running at 3.2 times last year's rate against user-facing changes at 1.36 times means each unit of shipped change now drags about 2.35 times the internal churn it did a year ago, a rise of roughly 135% [1]. The agents produced work, and most of the work landed on Meta's own platforms and infrastructure rather than in anything a user could open [3][4].
The incident line behaves the same way. Major technical and security incidents rose 40% year over year, while the hours staff burned putting them out rose 70% [5][6]. Divide those and each incident cost about 21% more time to close than an incident a year earlier [2]. Volume went up and so did difficulty, which is what the April internal post describing unchecked agents taking "large-scale, disruptive actions that humans are unlikely to execute" would predict [7]. When a person breaks something, there is a person who remembers doing it.
Worth noting where the 220% came from: Meta's own CTO, in an early June post [3]. The velocity number was not produced by a critic. Reuters reports that executives were separately seeing "reliability warning signs" caused by the coding surge [c7b]. Sanchit Vir Gogia of Greyhound Research puts the failure in sequencing rather than credulity, saying Meta "booked a forecast as capacity" and budgeted an improvement that had not yet arrived in production [8]. His order of operations is prove the action, then widen the authority, then remove the human control [9].
Terra Higginson of Info-Tech Research Group makes the measurement point plainly: output should not stand in as a proxy for productivity, and what she is seeing is action without the outcome anyone wanted [10]. Justin Greis of Acceligence says AI can make an organisation extraordinarily busy without making it more productive [11]. Then, in the same write-up of a productivity forecast that collapsed under audit, Conifers.ai's Tom Findling suggests telling the board that you may not get a 500% boost but IT can show a way to 300% [15]. The multiple changed. The habit of quoting one before the capability exists did not.
The employee-side detail is the one that will travel. After the April mandate to install tracking software on US employees' devices to capture keystrokes and mouse clicks, teaching agents to replicate how humans use a computer, staff rebelled on the reasonable suspicion that they were training their replacements [12]. Meta's remedy was more travel and social spending plus better snacks in the office microkitchens, and morale did not move [13].
One caveat sits under all of it. These are Reuters' readings of Meta's internal posts [2], and a company that halted a plan to cut up to 60% of some teams on the strength of its own dashboards has not shown anyone the dashboards [1].
Ranked by verification strength, evidence, and original report placement.
Earlier this year Meta was ready to cut up to 60% of the members of some teams and replace them with AI, as part of Project OT (Organization Transformation), an initiative to make Meta "AI native".
Meta backed off the plan at the last minute after internal data showed it was not working out, according to a Reuters investigation published Wednesday.
Code changes made to the internal software platforms and infrastructure that Meta employees used on the job were up 220% year over year, according to an early June post by Meta CTO Andrew Bosworth cited by Reuters.
Changes that led to new or upgraded features reaching Meta users were up only 36%.
Major technical and security incidents, such as service disruptions and possible data leaks, spiked 40% from the previous year.
The time staffers had to spend firefighting those incidents was up 70%.
Follow any of these and your For You feed starts watching them — no settings page required.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Specific internal metrics, but relayed second-hand and unverified
The core numbers are concrete and internally sourced (a CTO post, internal reliability posts, incident and firefighting deltas), which is stronger than vendor-marketing evidence. But this cluster contains a single trade-press summary of another outlet's investigation: no primary documents, no absolute baselines, no Meta response, and no independent replication. That caps evidence strength around the midpoint.
Agentic coding deployed at production scale internally; workforce substitution reversed
Adoption of AI coding agents inside Meta is real and large enough to move company-level metrics - a 220% rise in internal code changes, a 40% rise in major incidents, and a mandated telemetry-capture program to train agents on human computer use. What did not get adopted is the organizational step: the Project OT cuts of up to 60% on some teams were shelved. So tooling adoption is high while substitution adoption is near zero, and the score reflects that split.
AI-native replacement claims ran well ahead of measured outcomes
The gap is positive and sizeable: an 'AI native' plan to remove up to 60% of some teams was budgeted against capability that had not arrived, and the internal numbers showed activity rather than delivered outcomes - 220% more internal code changes for 36% more user-facing change, plus more incidents and more firefighting. Named analysts characterize this directly as booking a forecast as capacity and mistaking output for productivity. The gap is not larger because the reporting is corrective rather than promotional, and because the article itself documents the reversal rather than amplifying the original claim.
Corrective reporting, but all commentary comes from AI-governance sellers
Every interpretive voice in the cluster has a commercial stake in the lesson being drawn: two research firms (Greyhound, Info-Tech) and two consulting or product CEOs (Conifers.ai, Acceligence), one of whom supplies a board-pitch script contrasting a 500% promise with a 300% deliverable. The publisher is an enterprise IT trade outlet whose audience is the buyer of that advisory work, and no disclosure of these interests appears. The underlying facts originate from leaked internal Meta data via Reuters, which is an independent path, so incentive pressure sits above the midpoint rather than at the top.
Directionally credible, numerically fragile
Confidence is moderate. The direction of the story - agent output outpacing shipped value, incident load rising, headcount plan pulled - is coherent, internally sourced, and consistent across two separate metric families. But it is one publisher summarizing another outlet's investigation, with no Meta comment, no primary documents, no baselines, and derived ratios computed from percentage deltas alone. Individual figures should be treated as reported rather than established.
product
Meta ran pods-plus-agents for a year and shelved it. Its own scoreboard says why1 distinct publisher
build
Meta's agent plan produced 220 percent more code changes and 36 percent more shipped work3 distinct publishers
product
Meta owns the models and the data centres, and still pays Microsoft to rent someone else's1 distinct publisher
build
Meta moved Quest's front door: what the August 25 changes make VR teams re-test1 distinct publisher
Distinct publishers with included, body-backed reporting in this cluster.
1 article · August 26, 2026