Build1 publisher3 min readPublished
Agents stall on tool surfaces, not models: tool design as an engineering discipline
A dev.to series entry on the harness tool layer argues fifty bespoke tools are worse than read_file and bash. The usable test is whether one tool already expresses another.
The Engineer · Build desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened
- "Harness Engineering - Part 4: The Tool Layer" was published on dev.to as the fourth entry in a 10-part series described as a journey from raw language model to production-ready agentic system.
- Tools are the set of functions the harness exposes to the model; the harness tells the model the function names, the parameters they take, and what they do.
- The article states its subject is the design decisions that separate a tool surface a model can actually use from one that constantly frustrates it.
- On a given turn the model can emit a structured request to call a tool, for example {"tool": "read_file", "parameters": {"path": "/etc/hosts"}}.
- The harness sees the tool request in the model's response, runs the actual function, and feeds the result back into the model on the next loop iteration; the article describes this as two-way traffic: the model requests, the harness executes, the harness returns.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
Part 4 of the Harness Engineering series on dev.to takes up the tool layer: the set of functions a harness exposes to a model, along with their names, their parameters, and the descriptions of what they do [1][2]. The framing matters more than the mechanics, because the piece treats the tool surface as the thing that decides whether a capable model can act or spends its turns frustrated [3].
The mechanic itself is small. The model emits a structured request, something like a call to `read_file` with a path of `/etc/hosts` [4]. The harness sees the request, runs the real function, and returns the result into the next loop iteration [5]. That is the whole contract: model requests, harness executes, harness returns [5]. The article's analogy is the operating system offering system calls to a running program, with the model as the consumer instead of the process [6]. Without this, a loop is a hollow shell: it calls the model, the model produces text, and the text goes nowhere [7].
The load this layer carries is specific. Of the five gaps the series attributed to a raw model in Part 1, three close through tools: no execution, closed by tools that run code, touch files and hit APIs; no fresh knowledge, closed by search, retrieval and query tools; and no environment, since tools are how the model reaches whatever environment the harness sets up [8]. Persistence and verification are assigned to other components in later parts [9]. Every agent capability you have seen has arrived this way: an edit tool when Claude Code writes a file, a search tool for a research agent, a `get_order` tool for a support agent [10].
Which is why the failure mode named in the piece deserves to be treated as a defect rather than a style preference. Give every use case its own bespoke tool and you arrive at fifteen or fifty of them [11]: `read_python_file`, `read_javascript_file`, `read_config_file`, `list_files_in_directory`, `list_files_matching_pattern`, `run_python_script`, `run_shell_script`, and onward [12]. The article's position is that this surface is almost always worse than two well-named tools, `read_file(path)` and `bash(command)`, which express everything on the long list and a great deal it does not cover [13][14]. Seven entries collapse to two, and the seven were never more capable [15].
That gives a criterion an engineer can actually apply without a benchmark run: for each tool in your surface, check whether an existing tool already expresses it. If it does, the extra entry is description budget spent twice and one more branch for the model to disambiguate on every turn. Composability is testable in this narrow sense; "the model gets confused" is not.
Two honest limits. The article says good tool surfaces share three properties, but the text available here breaks off mid-sentence inside the first one [16], so the other two criteria are not yet on the table. And the author sells a Udemy course and a live Maven workshop alongside the series, described as optional [17].
Worth watching whether Part 5, on context engineering, connects tool descriptions to context budget, since that is where a bloated surface actually costs money [18]. Meanwhile, run the substitution test on your own surface and count how many entries survive it.