Build1 publisher3 min readPublished
Palo Alto Networks' Cortex XCOR keeps engineers approving what its root-cause agents recommend
Palo Alto Networks launched Cortex XCOR, built by its January Chronosphere acquisition, to root-cause outages with AI agents that default to human sign-off. For on-call teams that makes it a faster first investigator, with a person still approving each fix.
The Engineer · Build desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened
- When an alert fires, XCOR starts a specialized agent that reasons through the cause and recommends actions and mitigations, said Martin Mao, Palo Alto's observability chief.
- Other specialized agents mirror the remaining jobs an observability user does, such as tuning alerts and dashboards and optimizing data volumes.
- XCOR Operator, a conversational assistant, fronts those agents and is pitched as helping operations teams "match AI-coding velocity".
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- decision Widening autonomous permissions is the adoption decision that matters, because each step moves an action from a proposal a person approves to something the agent executes alone.
- constraint Because the agent starts only when an alert fires, its coverage ends where a team's alerting ends; an outage that never trips an alert gets no automatic investigation.
- exposure When reasoning models pick their own route to an answer, the access a team grants them, and no fixed workflow, sets what the agent can touch in production.
The New Stack's headline says XCOR traces outages in minutes and still pages engineers [4]. The report does not include a measured investigation time, an accuracy rate, or the list of actions that sit behind the autonomous permissions [15].
For a time-to-root-cause figure to transfer to another shop, the agent has to reach the telemetry an engineer would have opened. Palo Alto's own explanation for why this was not built sooner is about that data. It has said the scale and cost of cloud-native architectures stood in the way, and that pre-AI-boom automation was sluggish [8]. Chronosphere, the company it bought in January, sold an observability platform and a telemetry pipeline [6]. The XCOR team is the Chronosphere team [7]. I'd expect XCOR's results to track how much of a team's telemetry already passes through that pipeline.
On call, the change is in what the page asks of the engineer. The agent has started reasoning before anyone opens a laptop [9]. Under the human-in-the-loop default, a person still decides what happens to its recommendation [2]. Mao described that stage with an aviation comparison. "Right now we're at the moment where SREs will operate like airline pilots; they can rely on autopilot for smooth flying, but you still need an experienced pilot in the cockpit when something goes wrong," he said [3].
His stated goal goes further. "We don't want to keep waking engineers up in the middle of the night to help them make sense of dashboards; to be clear... we don't want to wake them up at all," Mao said [1]. The route there is the permissions setting. "While XCOR defaults to a human-in-the-loop model, engineering leaders can expand its autonomous permissions over time and finally get a full night's sleep," he said [2]. Read literally, the sentence has the leaders getting the sleep.
The design choice underneath shows up in another of his remarks. "We're now moving from strong, pre-defined specs for features and workflows to giving the reasoning models the right access and capabilities and allowing them to discover the various paths to reach an answer," Mao said [12]. The approach suits incidents nobody wrote a runbook for. It is also harder to review. An engineer approving a mitigation at night is checking a path the agent chose, with no fixed workflow to compare it against.
The pitch for XCOR Operator leans on code volume [10]. The New Stack cites a BairesDev Dev Barometer analysis in which 42% of developers say AI writes at least half their code, up from 12% a year earlier [13]. That is a rise of 30 percentage points [14]. The survey measures who writes the code. Incident volume is a separate number.
In my context, a team with a maintained alert set and runbooks it trusts, I'd run XCOR in recommend-only mode first. Each proposed root cause gets scored against what the on-call engineer actually found. Permissions get widened once that hit rate is known. Palo Alto ships the product in that human-approval mode by default [2].
What to watch
- Published time-to-root-cause or accuracy figures from a named XCOR customer, with the workload and alert volume described.
- Palo Alto documentation of which actions each XCOR autonomous permission level unlocks and how agent actions are logged for review.
- Whether XCOR's agents can investigate using telemetry that does not pass through the Chronosphere pipeline.