Product1 distinct publisher3 min readUpdated
A devops.com essay argues durability, per-step identity and decision-level observability belong in the agent runtime, not layered on later. The failure economics support the claim.
The Product Desk · Product desk
Compiled by The Product DeskSomething wrong?How this is made
A devops.com essay argues that the most expensive mistake teams are making with AI agents right now is treating them as microservices with a language model bolted on [1]. That matters because the microservices playbook allowed resilience to be added later, and the piece contends that option has quietly closed [7].
The argument rests on a property, not a preference. Microservices were hard to distribute but deterministic in behavior: a service took a call, returned a result, and when it failed it failed in milliseconds and threw an error you could see [3]. An agent workflow can run for hours or days, touch a dozen systems, and make a non-deterministic decision at step three that nobody discovers was wrong until step forty [4]. That is roughly 37 subsequent steps executed on a bad premise, with nothing red on a dashboard [15]. The essay's framing is blunt: a broken microservice pages you, while a misbehaving agent returns a confident wrong answer and continues [5].
The consequence is an inversion of the usual ops assumption. When an agent workflow fails five steps into a multi-day run, detecting that something broke is often harder than recovering once you know, and the resilience teams intended to add later never gets added [8]. So the recommendation is that recovery, per-step permissions and identity, and enough visibility to reconstruct why an autonomous system made a given decision all move into the foundation [6].
The more useful part of the piece is that existing durable execution engines are not the answer. Those engines were built for high concurrency and millisecond steps: payment flows, order processing, thousands of short workflows multiplexed per worker [9]. Agent workloads invert every one of those assumptions. A single step might be a 45-second model call, roughly 45,000 times the duration those systems were tuned for [10][16]. Another might pause for hours waiting on a human, during which the process should be able to die and resume cleanly, because nobody cares whether the resume takes seventeen milliseconds or two seconds [10]. And replay, the mechanism most durable engines depend on, does not hold: you cannot replay a non-deterministic step and expect the same result [11]. The conclusion is that the durability layer has to be rebuilt for a workload it was never designed to carry, not adopted off the shelf [7].
Two caveats. First, this is a single-sourced argument, and it lands on "agentic durable execution" as the emerging requirement: a platform layer that carries workflows to completion while producing trustworthy evidence of what happened, without sacrificing portability or open standards [13]. That is a category description, and category descriptions tend to arrive attached to products. Second, the excerpt's security section is titled "Security and governance need a new model, not a stronger boundary" and begins by noting that microservices pushed authentication to the edge, but the supplied text breaks off before the replacement model is described [14][17]. Per-step identity is asserted as a requirement [6]; the mechanism is not.
The testable part is the timeline. The author claims that treating agentic AI as just another microservice is an architectural bet coming due in the next 12 to 18 months [12]. Worth watching whether the first wave of production agent incidents shows up as detection failures rather than crashes, since that is the specific prediction here [8].
Follow any of these and your For You feed starts watching them — no settings page required.
Ranked by verification strength, evidence, and original report placement.
The durable execution engines built for the microservices era assumed high concurrency and millisecond steps, payment flows, order processing, and thousands of shorter workflows multiplexed per worker.
A single agent step might be a 45-second model call; another might pause for hours while it waits on a human, and during that wait the process should be able to die and resume cleanly later, because nobody cares whether the resume takes seventeen milliseconds or two seconds.
AI agents break key microservices assumptions: they can run for hours or days, make non-deterministic decisions, and continue operating after taking a wrong turn without ever producing a traditional error.
Microservices-era systems stayed simple in one crucial way: they were deterministic. A service received a call and returned a result, and when it failed, it failed in milliseconds and threw an error you could see.
An agent workflow can run for hours or days, touch a dozen systems, and make a non-deterministic decision at step three that you do not discover was wrong until step forty; nothing throws an error and nothing lights up red on a dashboard.
Where a broken microservice pages you, a misbehaving agent sends a confident, wrong result and moves on.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Single opinion essay, internally coherent, externally uncorroborated
One trade-press essay from one publisher carries the entire cluster. Its descriptive claims about microservices determinism, agent non-determinism and the mismatch between millisecond-step durable execution engines and 45-second or hours-long agent steps are specific and internally consistent, which is why evidence is not scored at floor. But there are no benchmarks, incident reports, deployments, named systems or third-party corroboration anywhere in the supplied material, the prescriptive and forecast claims rest on argument alone, and the body is truncated mid-word so disclosures are unavailable.
No adoption signal in supplied material
The supplied source reports no release, deployment, benchmark, pricing or usage disclosure, names no product or project implementing agentic durable execution, and gives no counts of teams following or rejecting the pattern. There is no basis to score adoption, and none should be inferred from the essay's assertion that a requirement is 'emerging'.
Prescriptions and timeline outrun the evidence offered
The descriptive middle of the essay is roughly aligned with what it demonstrates: the failure-economics reasoning (37 steps on a bad premise, non-replayable steps, timescale mismatch) is self-consistent and does not overreach. The overstatement sits at the edges, where a superlative ('the most expensive mistake'), a dated forecast (12-18 months) and a named platform category are asserted with no data, adoption or named systems behind them. Positive but moderate: the gap is one of unproven scope and timing, not of contradicted facts.
Category-defining trade-press essay, disclosure not visible
Observable from the supplied text: the piece is thought-leadership in a vendor-facing trade publication that concludes by naming a product category teams are told they need, specified in procurement-ready terms (completion guarantees, trustworthy evidence, portability, open standards). That structure aligns the argument with whoever sells such a layer. Scored mid-range rather than high because the supplied body is truncated before any byline, affiliation or disclosure, so no commercial relationship is established, and the technical critique stands on its own reasoning rather than on comparisons that favour a named product.
Low-moderate: one publisher, no adoption data, truncated text
Confidence in this assessment is limited by structural facts of the cluster rather than by ambiguity in what was said: a single publisher, a single item, zero adoption observations, and a body that ends mid-word. What the article claims is clear and quotable, so claim-level extraction is reliable; whether those claims are true, timely or commercially motivated cannot be triangulated within this material.
invest
The card networks just picked the referee for agent checkout, and it looks like EMVCo2 distinct publishers
build
An empty array is a claim about your query: verify identifiers before you trust the metric1 distinct publisher
product
OpenTelemetry is free; the collector fleet, the retention policy and the on-call rota are not1 distinct publisher
build
Per-developer environments hit their ceiling the day one engineer ran five agents1 distinct publisher
Distinct publishers with included, body-backed reporting in this cluster.
1 article · August 19, 2026