Build1 publisher3 min readPublished
The bug in your multi-agent system is not the model, it is the open HTTP request
A dev.to write-up argues that synchronous agent chains break on restarts, slow tools and approval waits, and that the fix is an event bus plus a policy layer before any real action.
The Engineer · Build desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction
What happened
- A post published on dev.to carries the headline "Event-driven AI agents: Build multi-agent workflows that survive production failures".
- The post argues that AI agents become fragile when connected as long synchronous chains, and that an event bus lets them work independently, wait for people and tools, recover after restarts, and place policy between a model's recommendation and a real action.
- Most AI agent demos fit inside a single request: User -> Agent -> Tool -> Agent -> Response. The agent makes a plan, calls a tool, gets an answer, and returns a response.
- A real workflow needs data from several systems; one tool takes five minutes; another agent has to review the result; a production change needs human approval; and while the approval is pending, one of the services restarts.
- The post states: "The model is no longer the hardest part. Coordination is."
Compiled by The EngineerSomething wrong?How this is made
Why it matters
A post published on dev.to under the headline "Event-driven AI agents: Build multi-agent workflows that survive production failures" argues that once agents meet real workflows, the model stops being the hard part and coordination becomes the hard part [26][4]. Its practical claim is narrower and more useful than the usual agent discourse: put an event bus between agents so they can wait for people and tools and recover after restarts, and put a policy check between a model's recommendation and a real action [1].
The demo shape is familiar: user, agent, tool, agent, response, all inside a single request [2]. The production shape is not. The post's example workflow needs data from several systems, includes one tool that takes five minutes, requires another agent to review the result, requires human approval for a production change, and then has a service restart while that approval is pending [3]. No prompt fixes that. What the agents need is the ability to run independently, survive failures, and resume after the original request has ended [5].
The failure list is the part worth pinning to a wall. A timeout can leave a tool still running after its caller gave up; a retry can execute the same action twice; a restart can erase the current plan; and adding a security review means changing an integration that already worked [8]. Human approval exposes it most cleanly, because a decision can take minutes or hours, and holding an HTTP request open across that wait is a poor way to preserve a workflow [9].
The alternative is that an agent publishes what happened and other components decide whether that fact matters to them [10]. The worked incident runs ServiceLatencyIncreased, DiagnosisRequested, DependencyFailureSuspected, RecoveryProposed, HumanApprovalRequired, RecoveryApproved, RecoveryCompleted [11] - seven events for the seven steps the operations agent has to perform, from collecting metrics to verifying recovery [6][12]. A service health agent can request a diagnosis without knowing which agent will take it, and the diagnostic agent can report a suspected dependency failure without calling remediation directly [13]. An audit service, an observability pipeline and a security agent can all react to the same event, and a consumer that is down can process it after it recovers, subject to the broker's retention settings [14]. The post notes this is the same producer and consumer separation described in AWS Prescriptive Guidance for event-driven AI, and is not tied to any cloud or broker [15].
The other half is message taxonomy, because systems get dangerous when every message is treated as interchangeable [16]. An event records a fact, such as checkout-api at 1840ms p95 [17]. A command requests an action, such as ScaleService to 12 replicas, carrying an idempotency key like "incident-204:scale:12" [19]. An agent decision is a recommendation, with evidence and confidence, and still only a proposal [20]. An LLM saying "scale the service" means neither that the service was scaled nor that the action was authorized [21].
So the chain becomes event, agent decision, proposed command, policy check, authorized command, execution, outcome event [22]. Policy checks permissions, limits and approval requirements; deterministic code makes the change; the executor publishes what actually happened [23]. The payoff is that a model can get more capable without quietly acquiring permission to restart production or issue a refund [24].
Two things to watch in your own stack. First, whether the broker's retention window is actually longer than your slowest human approver, since that is the setting the recovery story depends on [14][9]. Second, whether execution state is being stored somewhere other than conversational memory; the post treats those as solving different problems [25], and a restart that erases a plan is the test. If events cross team or platform boundaries, the CloudEvents specification is the vendor-neutral envelope the post points at for metadata [18].