Product1 publisher3 min readPublished
StackGen wraps production agents in one shared record and a logged policy harness
Autonomous Operations Factory puts four operations agents behind a single record of the environment and one audit log. The urgency numbers behind the pitch come from StackGen's own reliability report.
The Product Desk · Product desk

What happened
- StackGen launched the Autonomous Operations Factory, a governed layer where specialized agents share context and split provisioning, deployment and incident response, built on a system it calls Aiden OS.
- Aiden OS pairs a shared record of the environment with a harness that enforces policy and logs what every agent does.
- Four agents cover infrastructure operations, DevOps, site reliability engineering and observability, and customers can bring their own agents onto the factory to inherit the same guardrails and audit trail.
- StackGen says it has documented at least nine cases since last year in which an agent took destructive action against a live production system on its own.
- The factory is in preview for Amazon Web Services, Microsoft Azure, Google Cloud and Oracle Cloud users, with all four agents available.
Compiled by The Product DeskSomething wrong?How this is made
Why it matters
- decision Teams that have bought one agent per operations function now choose between adding another point agent and re-basing the ones they already run on a shared record.
- exposure Whoever signs off on an agent acting in production without a human approval owns the after-action reconstruction, and a per-agent log is what that person gets asked to produce.
- constraint Adopting the factory puts the authoritative account of what is deployed and what changed inside a vendor's world model, and the agents a team already operates take their permissions from there.
- cost StackGen did not state a price for the paid pieces, so the reliability agent's free edition is the only part a team can put in front of its own incidents without a procurement conversation.
StackGen's picture of the failure is small and recognizable. A delivery pipeline repairs a failed build with no knowledge that the error budget is already spent [6]. Operations work still sits in four places on four sets of tooling, the company argues, and a point agent bolted onto one of them makes that one faster without connecting it to the others [5][7].
What the factory changes is where the record lives. Agents read one model of what is deployed, what changed, what broke and what fixed it, and they pass through one set of policy checks and approval flows [3]. A reliability investigation then opens with deployment history already loaded, and a deployment is gated against live infrastructure drift before it ships, according to StackGen. Anything one agent learns becomes available to the others [11].
The urgency numbers are the company's own. Its State of Reliability 2026 report attributes roughly 10 percent of disclosed outages this year to AI and calls that a six-fold rise over three years [8]. Divide 10 by 6 and the base three years ago was about 1.7 percent [9]. The share is small and rising quickly, and StackGen compiled the count.
Outside confirmation is thinner. Dan Twing, an analyst at Enterprise Management Associates, said the approach is sound in principle, that agents working in production need "lifecycle context and governed authority" rather than telemetry alone, and that connecting infrastructure, delivery, observability and reliability work through a shared model gives StackGen "a credible foundation for safer, more coordinated autonomous production operations" [16]. The one named customer is partway in. OneTrust automates observability and incident response with StackGen today and is working toward automating the rest of the operations life cycle, DV Lamba, its chief product and technology officer, said [15].
"Writing code with AI is faster than ever; running what it produces is not, and that gap is where enterprises lose money and take on more risk," said Sachin Aggarwal, co-founder and chief executive of the San Francisco company, which was called appCD until it rebranded in 2024 alongside a $12.3 million seed round [13][18].
The release also ships the Autonomy Index, benchmarks inside the product that score how autonomous each stage of the factory is, by team and by application type [12]. Autonomy measures how much a team has handed over. A score that climbs while incidents hold steady records more delegation and no improvement. The index sits beside a team's recovery-time numbers; it does not substitute for them.
Two facts about a team's own setup sort the decision. The first is whether any agent can take an action in production without a human approving it. The second is whether that action could be reconstructed afterward from logs the team already keeps. Where the answer to the first is no, the purchase is a new record of the environment, and the governance is a feature waiting for a use [3]. Where the first is yes and the second is no, the harness is the product, and the thing to test during the preview is whether it stops an action the team would have stopped itself.
What to watch
- Whether any independent reliability dataset reproduces the 10 percent AI-attributed outage share or the nine destructive-agent cases StackGen counted.
- Pricing for the paid agents when the four-cloud preview turns into general availability.
- Whether OneTrust extends automation past observability and incident response into the rest of its operations life cycle.