Product1 publisher3 min readPublished
AWS ships CloudWatch Omni with 17 evaluators that grade AI agents' answers
AWS made CloudWatch Omni generally available with 17 built-in evaluators that score an AI agent's answers on live production traffic. For operators, the job moves from confirming an agent is running to deciding whether an automated score is good enough to gate a release.
The Product Desk · Product desk

What happened
- Agent traces share one CloudWatch data store with application and infrastructure telemetry, so an investigation can follow a bad tool result down to an exhausted database connection pool.
- Omni is built on OpenTelemetry and runs outside the AWS Management Console, as a standalone web app that signs users in through Okta or Microsoft Entra ID.
- An extension for Visual Studio Code, Cursor and Kiro shows traces while developers run agents locally, and it needs no AWS account.
- AWS cites an IDC forecast of more than 1 billion deployed agents by 2029, a volume the SiliconANGLE column says no operations team could review by hand.
Compiled by The Product DeskSomething wrong?How this is made
Why it matters
- decision Once evaluators run on live traffic, someone has to pick the score drop that triggers an alert and own the page when it fires.
- capability Regression tests can be built from real production traffic, so a new prompt version is checked against the requests customers actually sent.
- exposure Application owners and AI engineers, the people AWS says its console was not built for, now hold the traces and scores that show a wrong agent answer, and with them the accountability for it.
The agent met its latency target and threw no errors. It also called the wrong tool and gave the customer a wrong answer. SiliconANGLE's column on Omni opens with that failure to argue that uptime metrics miss what matters for agents [17]. AWS's answer is to put answer scoring into the product it positions as the next generation of CloudWatch [2].
Both customers named at launch are large operations. "At Sony, our enterprise-wide agentic AI platform now supports hundreds of proof-of-concept and production workloads," said Masahiro Oba, senior general manager of the AI Acceleration Division at Sony [8]. "With Amazon CloudWatch Omni, I can go from a single trace directly to evaluation, AI analysis, comparison or dataset creation," he said [9]. Capital One was a design partner. "Capital One operates one of the largest observability footprints in financial services," said Parvez Naqvi, the bank's managing vice president of cloud platform and resilience engineering [16].
Teams will tell themselves that a rising helpfulness score means customers are getting help. What the product actually does is attach evaluator scores, for measures such as faithfulness and routing correctness, to captured traces [5]. The column says latency and error rates "can't tell you whether an answer was right, but scoring can" [18]. That holds only when the score agrees with a person reading the same trace. The account does not include pricing or any comparison of evaluator scores with human review.
The OpenTelemetry base means instrumentation written for Omni can be pointed at other backends [3]. The investigation across layers is tied to CloudWatch. It works when the API and the database behind the agent also report there [14], and AWS DevOps Agent is switched on by default in those investigation sessions [15]. The column notes that many startups can already trace large language model calls [19].
I'd sort the decision on two questions. How many agents does the company run, and does its application and infrastructure telemetry already live in CloudWatch? Many agents with telemetry in CloudWatch is where every part of the pitch applies. With many agents and telemetry in another vendor's tool, the evaluators still work on agent traces, but the cross-layer investigation does not. A handful of agents on AWS justifies a trial of the editor extension and a few evaluators on a single agent. For a small team whose telemetry sits elsewhere, Omni adds little beyond what a hand-checked test set already provides.
Whatever the quadrant, one check has to come before any evaluator score gates a release or pages someone. Run the evaluator against a sample of production traces that a person has already marked right or wrong. If the two disagree often, keep the score on a dashboard and out of the release process.
What to watch
- Whether AWS publishes pricing for continuous evaluation on live traffic, which decides whether scoring every production trace is affordable.
- Any published comparison of Omni's built-in evaluator scores against human reviewers' judgements on the same traces.
- Whether customers beyond Sony and Capital One report running Omni investigations across services that export telemetry from outside AWS.