Skip to content

Product1 publisher3 min readPublished

Agents Are Not Microservices With an LLM Attached, and the Retrofit Never Arrives

A devops.com essay argues durability, per-step identity and decision-level observability belong in the agent runtime, not layered on later. The failure economics support the claim.

The Product Desk · Product desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened

  • The article states that the most expensive mistake teams are making with AI agents right now is treating them as microservices with a language model bolted on.
  • AI agents break key microservices assumptions: they can run for hours or days, make non-deterministic decisions, and continue operating after taking a wrong turn without ever producing a traditional error.
  • Microservices-era systems stayed simple in one crucial way: they were deterministic. A service received a call and returned a result, and when it failed, it failed in milliseconds and threw an error you could see.
  • An agent workflow can run for hours or days, touch a dozen systems, and make a non-deterministic decision at step three that you do not discover was wrong until step forty; nothing throws an error and nothing lights up red on a dashboard.
  • Where a broken microservice pages you, a misbehaving agent sends a confident, wrong result and moves on.

Compiled by The Product DeskSomething wrong?How this is made

Why it matters

A devops.com essay argues that the most expensive mistake teams are making with AI agents right now is treating them as microservices with a language model bolted on [1]. That matters because the microservices playbook allowed resilience to be added later, and the piece contends that option has quietly closed [7].

The argument rests on a property, not a preference. Microservices were hard to distribute but deterministic in behavior: a service took a call, returned a result, and when it failed it failed in milliseconds and threw an error you could see [3]. An agent workflow can run for hours or days, touch a dozen systems, and make a non-deterministic decision at step three that nobody discovers was wrong until step forty [4]. That is roughly 37 subsequent steps executed on a bad premise, with nothing red on a dashboard [15]. The essay's framing is blunt: a broken microservice pages you, while a misbehaving agent returns a confident wrong answer and continues [5].

The consequence is an inversion of the usual ops assumption. When an agent workflow fails five steps into a multi-day run, detecting that something broke is often harder than recovering once you know, and the resilience teams intended to add later never gets added [8]. So the recommendation is that recovery, per-step permissions and identity, and enough visibility to reconstruct why an autonomous system made a given decision all move into the foundation [6].

The more useful part of the piece is that existing durable execution engines are not the answer. Those engines were built for high concurrency and millisecond steps: payment flows, order processing, thousands of short workflows multiplexed per worker [9]. Agent workloads invert every one of those assumptions. A single step might be a 45-second model call, roughly 45,000 times the duration those systems were tuned for [10][16]. Another might pause for hours waiting on a human, during which the process should be able to die and resume cleanly, because nobody cares whether the resume takes seventeen milliseconds or two seconds [10]. And replay, the mechanism most durable engines depend on, does not hold: you cannot replay a non-deterministic step and expect the same result [11]. The conclusion is that the durability layer has to be rebuilt for a workload it was never designed to carry, not adopted off the shelf [7].

Two caveats. First, this is a single-sourced argument, and it lands on "agentic durable execution" as the emerging requirement: a platform layer that carries workflows to completion while producing trustworthy evidence of what happened, without sacrificing portability or open standards [13]. That is a category description, and category descriptions tend to arrive attached to products. Second, the excerpt's security section is titled "Security and governance need a new model, not a stronger boundary" and begins by noting that microservices pushed authentication to the edge, but the supplied text breaks off before the replacement model is described [14][17]. Per-step identity is asserted as a requirement [6]; the mechanism is not.

The testable part is the timeline. The author claims that treating agentic AI as just another microservice is an architectural bet coming due in the next 12 to 18 months [12]. Worth watching whether the first wave of production agent incidents shows up as detection failures rather than crashes, since that is the specific prediction here [8].

Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories