Build1 publisher3 min readPublished
The payload is rebuilt every turn, so stop treating your prompt as a shipped artifact
A dev.to series argues "prompt engineering" is too narrow for agents: the package sent to the model is reassembled on every call, which makes it a runtime concern, not a file.
The Engineer · Build desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction
What happened
- dev.to published "Harness Engineering - Part 5: Context Engineering", part of a 10-part series described as a journey from raw language model to production-ready agentic system.
- The article defines the Context as everything fed into the model on a given API call.
- The article states that on any single turn the payload sent to the model typically includes the system prompt, the conversation history, any retrieved documents relevant to the current turn, the tool definitions, prior tool results, any attached files or images, and anything else the harness thinks the model needs to know right now.
- The article says some readers know this territory as prompt engineering, that the name is not wrong but narrow, because a prompt sounds like something you write once and ship while the reality of running an agent is that the payload changes every turn; hence the term context engineering.
- The article states that the whole package goes over the wire, the model reads all of it and produces one response, and then on the next iteration of the loop the harness assembles a new package, probably with the previous response added, maybe with new tool results appended, maybe with different retrieved content.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
Part 5 of the Harness Engineering series on dev.to makes a narrow but load-bearing argument: "prompt engineering" is not wrong, it is just too small a name, because a prompt sounds like something you write once and ship while the payload an agent sends changes on every turn [1][4]. The consequence for anyone building one is that the thing you maintain is not a string in a repo but an assembly step that runs before every model call [5].
The author defines the Context as everything fed into the model on a given API call [2]. Concretely, that payload typically carries the system prompt, the conversation history, any retrieved documents, the tool definitions, prior tool results, attached files or images, and anything else the harness decides the model needs right now [3]. That is six named categories plus an open-ended slot [7]. All of it goes over the wire, the model reads all of it, produces one response, and then the harness assembles a new package for the next iteration of the loop, probably with the previous response added, maybe with new tool results, maybe with different retrieved content [5].
The mechanism underneath is unglamorous: the model is stateless, every API call is independent on the model's side, and nothing persists between calls unless something outside the model puts it there [6]. You cannot tell the model to remember what was said five minutes ago or to reference a file discussed earlier; every fact, message, result and document has to physically be in the payload for the current call or it does not exist [8]. Continuity is a property of your harness, not of the model.
Read as a systems statement rather than a naming quibble, this is where the engineering shows up. The author calls context engineering arguably the deepest discipline in the whole harness, and describes the actual work as deciding what to include, what to compress, what to leave out, and when [9]. Compression and omission only make sense against a finite budget, which is the practical shift: you are sizing and prioritising a per-call payload, not editing copy. Of the three moving pieces the piece names -- system prompt, history, retrieved knowledge [10] -- only the system prompt is described as staying roughly the same throughout the conversation, prepended to every turn to tell the model who it is and what tools it has [11]. So the stable, reviewable, version-controllable layer is one slot out of the several the payload carries [12]. The rest is decided at runtime, per turn, by code you wrote and probably do not inspect.
Two caveats. The excerpt supplied breaks off mid-sentence in the system prompt section, so the treatment of history and retrieval is not visible here [13]. And the author sells a Udemy course and a live Maven workshop on building a harness, both described as optional alongside the free series [14].
What to watch is whether the remaining parts get specific where this one is conceptual: the series lists the filesystem and environment, memory, observability, the overall architecture, and a teardown of Claude Code still to come [15]. Memory and observability are where a per-turn payload budget becomes measurable rather than rhetorical: eviction and compaction policy, and the traces that tell you what was actually in the context when the agent got it wrong.