Skip to content

Build1 publisher3 min readPublished

Agent reliability is a harness problem, not a prompt problem

A dev.to write-up splits agent work into prompts, context and harness. The interesting part is the harness: tool execution, permissions, validation and recovery, all of it code you own.

The Engineer · Build desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened

  • A dev.to write-up distinguishes three layers: prompt engineering tells the model what to do, context engineering gives it the right information, and harness engineering builds the system that helps it act, verify and recover.
  • Early LLM application work revolved around prompts: improving instructions, adding examples and adjusting wording; that approach works well for simple tasks.
  • The article's worked example is a coding agent asked to add rate limiting to an existing Node.js API without breaking authentication.
  • For that task the article says a useful agent may need to understand the repository architecture, inspect authentication middleware, read engineering conventions, modify files, run tests, execute ESLint and TypeScript, inspect failures, correct its implementation, and possibly ask for human approval before changing sensitive infrastructure.
  • Anthropic describes context engineering as curating and maintaining the optimal information supplied to an LLM during inference, including conversation history, retrieved documents, tools, MCP resources, previous tool results, memory, application state and external data.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

A developer write-up on dev.to argues that reliable coding agents are built in three layers rather than one: prompt engineering covers how instructions are expressed, context engineering covers what information the model has at a given moment, and harness engineering covers the system that lets an agent act, verify and recover [1]. The consequence for anyone shipping this stuff is that most of the reliability work lives in code you own and can test, not in a text box you keep rewording.

The example the piece uses is ordinary enough to be useful: add rate limiting to an existing Node.js API without breaking authentication [3]. To do that, the write-up notes, an agent has to understand the repository architecture, inspect the authentication middleware, read the engineering conventions, modify files, run the tests, execute ESLint and TypeScript, inspect the failures, correct its own implementation, and in some cases stop and ask a human before touching sensitive infrastructure [4]. Almost none of that is a wording problem. It is tool calls, exit codes and permission checks. The early habit of improving instructions, adding a couple of examples and hoping for better behaviour holds up for simple tasks and stops there [2].

The middle layer has a known trap. Anthropic describes context engineering as curating and maintaining the optimal information supplied to a model during inference, spanning conversation history, retrieved documents, tools, MCP resources, previous tool results, memory, application state and external data [5]. Anthropic also warns that performance can degrade as more information competes for attention, which makes context a finite resource to be managed [6]. So the naive path of loading the whole repository, sending 200,000 tokens and asking the model to sort it out is not a shortcut [15]; the actual work is deciding what earns a place in the window [7]. Retrieval, memory, repository search, context compression, tool-result filtering and just-in-time retrieval all exist to serve that decision [8].

The harness is where the argument earns its keep. The mental model offered is Agent = Model + Harness [9], and LangChain defines the harness broadly as the code, configuration, tools, infrastructure, state and orchestration surrounding the model [10]. The architecture sketch names seven responsibilities: context management, tool execution, memory and state, permissions, validation, retry and recovery, and observability [11]. Six of those seven concern acting and controlling rather than what the model reads [16].

The test loop shows why that matters. A model can decide it should run the tests, but something outside the model must expose the testing tool, execute the command, capture stdout and stderr, enforce timeouts, return the relevant output, prevent unsafe commands, and hand control back so the agent can choose what to do next [12]. Each of those is a place where an agent quietly stops being trustworthy: swallowed stderr, no timeout, an approval gate that defaults to yes.

The layers are not rivals. Martin Fowler's discussion of coding-agent harnesses treats context engineering as one mechanism through which guides and feedback reach the agent [13], and the dev.to piece frames the three as stacked questions rather than competing approaches [14]. Worth noting what the material does not contain: definitions, examples and diagrams, but no measured error rates or before-and-after comparisons [17]. Treat it as a claim about where to spend engineering time, not as evidence of a specific reliability gain.

Watch whether harness configuration gets versioned and reviewed like application code, whether permission gates default to allow, whether test and lint failures are returned to the model in full or truncated to a useless tail, and whether context is budgeted and monitored the way any other finite resource is [6][11][12].

Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories