Skip to content

Build1 publisher3 min readPublished

Elastic's IT team says AI ROI has to be a query, and it instrumented every event to get one

A $2.5M time-recovery figure and a 30% self-resolution rate rest entirely on per-event telemetry captured in the first sprint. Elastic's own survey says 8% of teams have that.

The Engineer · Build desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened

  • Over six months, the Elastic IT team ran internal AI applications that returned $2.5 million in operational time to the business.
  • A conversational support assistant moved Elastic from zero digital resolution, where anything complex became a ticket, to 30% of support interactions closing without one.
  • Elastic says it can put those numbers in front of a finance team because they were measured from day one at the level of individual usage events.
  • For each use case, Elastic assigned a conservative time-saving goal and validated it with the teams doing the work.
  • Elastic gives the example that a support case summary saves about five minutes.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

Elastic's internal IT team has published an account of six months running its own AI applications, reporting $2.5 million in operational time returned to the business and a conversational support assistant that took the organisation from zero digital resolution to 30% of support interactions closing without a ticket [1][2]. The claim worth arguing about is not the dollar figure but the precondition attached to it: according to Elastic, those numbers exist only because usage was measured from day one at the level of individual events [3].

The mechanism is deliberately unglamorous. Each use case got a conservative time-saving goal, validated with the teams actually doing the work [4]. A support case summary was assigned roughly five minutes [5]. Multiply events by minutes saved by a standard labour burden and the return becomes a live KPI rather than an assumption that generative AI must be valuable because it is generative AI [6][7]. Elastic's framing is that ROI should be a query, not an anecdote [8].

Read the $2.5 million for what it is. It is recovered time, not recovered cash: hours went back to support engineers who had been searching for answers, and into roadmap work [9]. Whether that converts to money depends on what the roadmap produces, and the inputs are Elastic's own conservative estimates. The company states the formula but, as published, not the burden rate or the event volumes, so an outsider cannot recompute the total [10]. Elastic also sells observability tooling, which is the obvious reason to treat the write-up as a method to copy rather than a benchmark to cite.

The method is rarer than the intent. In Elastic's Landscape of Observability survey of 500 IT decision-makers, 85% said they planned to enable observability for their LLM applications and 8% had actually done it, a gap of 77 percentage points [11][12][13]. Separately, 93% report financial and business impact to leadership in some form, but only 19% do so regularly as part of an established process, leaving 74 points' worth of organisations that report occasionally or when asked [14][15][16]. Elastic calls the deferral measurement debt, on the grounds that a capability shipped without instrumentation borrows against a future demand to prove it was worth shipping [17].

The operational argument is stronger than the accounting one. A generative application can be fully available, comfortably fast, and completely wrong, so uptime and latency do not cover it [18]. Elastic tracks token consumption as the cost meter, retrieval quality because a grounded answer depends on what was retrieved, and use-case classification because knowing what people actually ask is the fastest route to a correction [19]. Launch instrumentation only tells you the application is running; the question that follows every model swap, prompt rewrite, or retrieval change is whether this version is better than the last, and that comparison requires a record of what came before [20]. A test suite answers only for the cases someone thought to capture, while live questions drift [21]. The baseline has to exist before you know you need it [22].

That is also why phase-two telemetry frequently is not recoverable: data ages out, and sampling discards whatever was never flagged as important [23].

Watch whether the 8% figure moves in the next survey round, and whether anyone publishes a burden rate and event count alongside a headline savings number. Also watch the naming problem Elastic raises: hand-built AI applications force someone to name every emitted field, including token counts, model, tool calls, and retrieval steps, which is where cross-team comparison usually dies [24].

Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories