Skip to content

Build1 publisher3 min readPublished

Your agent didn't misunderstand the prompt. It ran the wrong branch.

A dev.to essay splits agent architecture into five control layers and shows that only one of four failures in its worked example is fixable by editing instructions.

The Engineer · Build desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened

  • The dev.to piece asserts there are five control layers between a raw model call and a system you can trust with a business outcome: prompt, context, harness, loop, and graph.
  • The piece states that most teams staff and instrument only the first one or two of the five layers.
  • The failures that show up in production, per the piece, are wrong tool called, same mistake retried forever, and output routed to the wrong reviewer, and they live almost entirely in the layers nobody named.
  • Graph engineering is described as the newest and least understood of the five layers: it decides which component runs next, when agents work in parallel versus in sequence, and where a human has to sign off before anything expensive or irreversible happens.
  • The piece defines: MODEL CALL = prompt + context; AGENT = model call + harness + loop; SYSTEM = agents + deterministic steps + humans, connected by a graph; EVALS = evidence that every layer actually works.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

An essay published on dev.to argues that the agent failures operators keep escalating as prompt problems are structural, and it names the missing pieces: five control layers stand between a raw model call and a system you can trust with a business outcome, being prompt, context, harness, loop, and graph, and most teams staff and instrument only the first one or two [1][2]. That matters because the failure modes that generate actual production incidents, according to the piece, are wrong tool called, the same mistake retried forever, and output routed to the wrong reviewer, and those live almost entirely in the layers nobody named [3].

The decomposition is blunt: a model call is prompt plus context, an agent is a model call plus harness plus loop, a system is agents plus deterministic steps plus humans connected by a graph, and evals are the evidence that each layer works [5]. The graph layer is the one the author calls newest and least understood: it decides which component runs next, when agents run in parallel versus in sequence, and where a human has to sign off before anything expensive or irreversible happens [4].

The worked example is the useful part. A coding agent scoped to low-risk defects in an internal payments service passes on a clean sample repository and then fails four separate ways on the real one: it misses an architecture decision buried in the docs (context), runs a shell command with broader scope than intended (harness), retries the same failing test without changing its hypothesis (loop), and sends the pull request down the wrong review path (graph) [8][9]. The piece states that only one of those four traces back to the instruction layer, and it is not the one that caused the damage [10]. Three of the four sit outside the prompt entirely [14].

Note what the example prompt already said: stop and ask for approval if the fix changes an external contract [7]. The instruction was there. The pull request still went down the wrong review path [9]. An approval sentence in a system prompt is a preference, not a gate; the author's point is that the prompt cannot supply a missing design document, restrict a dangerous tool, or decide who reviews the output [11]. The harness is where the broad shell command should have been stopped, since it owns tools, file and shell access, sandboxing, permissions, timeouts, logging, and approval boundaries [12][13]. And MCP does not rescue this: the piece is explicit that MCP standardizes how an agent connects to tools but does not decide that an agent deserves production write access, which remains a matter of identity, least privilege, and approval policy in the host and its control plane [6].

The framing to keep is that these are concentric controls rather than pipeline stages, all five active at once, with the weakest layer setting the ceiling on reliability no matter how good the other four are [15][16]. Rewording a prompt may paper over one test case; fixing context assembly fixes the class [17].

Worth watching: whether teams start attributing incidents by layer in postmortems rather than by prompt diff, and whether eval suites get written per layer as the piece suggests, instead of only end to end [18].

Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories