Skip to content

Product1 publisher3 min readPublished

Dynatrace buys Arize AI to put agent evaluation next to its application traces

Dynatrace is buying Arize's tracing and response-quality evaluation for AI agents. Both product chiefs said their own customers had asked for the other side's telemetry, on a podcast where an analyst put most organizations at six to 15 observability tools.

The Product Desk · Product desk

Illustration accompanying Dynatrace buys Arize AI to put agent evaluation next to its application traces

What happened

  • Dynatrace is acquiring Arize AI, folding AI observability, evaluation and agent monitoring into its broader application observability platform.
  • The stated reason is that AI applications and agents can produce different outputs from similar inputs, unlike the deterministic software that logs, metrics and traces were built to monitor.
  • Dhinakaran said Arize customers had been asking for tighter links to production application telemetry while Dynatrace customers were asking for deeper AI evaluation.
  • Analyst Paul Nashawaty cited research on the same podcast putting 75% of organizations at between six and 15 tools for observability.

Compiled by The Product DeskSomething wrong?How this is made

Why it matters

  • decision Teams that bought a standalone LLM evaluation product now have a renewal conversation in which their observability incumbent can offer the same capability on an existing contract.
  • constraint Each new AI monitoring layer widens the set of places an on-call engineer looks before the first hypothesis, and a merger only shortens that search for failures that genuinely cross the eval-to-infrastructure boundary.
  • capability If a quality score and an infrastructure trace share one context, an automated remediation step can be handed both without a person correlating two products by timestamp.
  • precedent An observability platform buying an evaluation vendor makes evaluation look like a platform feature, and remaining standalone eval vendors will be quoted against a bundled price.

The on-call engineer paged for a fluent wrong answer can read, in the eval trace, that the response scored badly on quality. Finding out why means opening something else, because the agent was calling APIs and hitting databases that sit in telemetry the platform team owns [13]. "The agent systems and the software systems are joined at the hip," Dhinakaran said [9].

She was speaking on theCUBE Research's AppDevANGLE podcast, where analyst Paul Nashawaty interviewed her alongside Dynatrace chief product officer Steve Tack [12]. "Evaluating no longer just becomes about is it right or wrong," Dhinakaran said. "It becomes about actually measuring the quality of the responses, which is just a very fundamentally different problem" [5]. Tack took the wider view: "AI brings new problems, new domains to the space," he said [4].

The pitch around all of this is observability moving past detection, with enterprises expecting the platform to supply enough context for humans and AI agents to diagnose, remediate and potentially act on a problem [3]. What was actually bought is narrower: eval traces and application traces under one vendor. A shop at the top of the band Nashawaty cited adds a sixteenth tool the day it buys a standalone eval product, and one at the bottom adds a seventh [16]. The other quarter of organizations sits outside the band in one direction or the other [17].

The record supporting the thesis is thin. It is a conversation between the two product chiefs and one analyst. The largest usage figure in it belongs to Phoenix, Arize's open-source platform, which Dhinakaran said more than 4,000 enterprises use [6]. The commercial side is Arize AX, a managed environment for teams running AI systems at production scale [7]. The account does not state deal terms, an integration timeline, or how many Phoenix users pay for AX [15].

A test on your own incident log will settle the consolidation question faster than the deck will. Take the last five agent incidents and mark each on two axes. Axis one: whether the diagnosis was available inside the eval traces, or whether someone had to open a second product to find the slow tool call. Axis two: whether a person applied the fix, or a runbook did. Incidents where the diagnosis crossed products and the fix is automated are the ones a merged platform pays for, since the automated step needs the quality score and the infrastructure context in the same place. Incidents that stayed inside the eval tool with a human fixing them do not justify a bundle, and the separate vendor can keep competing on evaluation quality.

Tack argued that combining application and AI observability gives organizations a more complete system-level view [14]. Buyers can check that against a count of how many of their own agent incidents needed two consoles.

What to watch

  • Whether Dynatrace sells Arize AX as its own SKU or prices evaluation into existing platform licensing.
  • Whether Phoenix stays open source and usable without a Dynatrace contract after the deal closes.
  • Any published date for shared trace context between Arize eval traces and Dynatrace application telemetry.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories