Build1 publisher3 min readPublished
117 identical errors, zero bugs: when the defect lives in the orchestration
An overnight agent pipeline logged the same file-not-found error 117 times in seven weeks. Tracing it found no broken code, and the run reports never mentioned it once.
The Engineer · Build desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened
- Between May 15 and July 2 of this year, the session transcripts of the author's content pipeline accumulated at least 117 copies of the same error, 'File does not exist'.
- One error class, 117 occurrences, spread across seven weeks of overnight runs.
- When the author finally traced the error, they found no bug: not one line of code was doing anything other than what it was written to do.
- The author is an IT analyst, not a software engineer by training, and runs the content pipeline behind bestaiweb.ai with a colleague who is the actual programmer; the author's side is orchestration, audits and review.
- The author says the thing that cracked the case open was not a debugger, it was counting.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
An IT analyst who runs the overnight content pipeline behind bestaiweb.ai sat down with seven weeks of session transcripts and counted at least 117 copies of a single error, "File does not exist", logged between May 15 and July 2 [1][2][4]. Then came the tracing, and there was no bug: not one line of code was doing anything other than what it was written to do [3].
The shape of the pipeline explains where that leaves the defect. One topic consumes roughly 18 agent sessions across research agents, an article writer, a claim verifier, image generation and validators, all coordinated through agent orchestration [6]. Each agent is handed a small YAML brief that names the article, the fact sheet and the output location; the brief is written by deterministic TypeScript and Python and read by an LLM [7]. The author puts the whole story at that handoff: code on one side of the file, an interpreter on the other [23].
The error text is the tell. Verbatim, it reads "File does not exist. Note: your current working directory is /Users/userxy/code/your-project." [9]. According to the author, the helpful note is the trap, because it supplies the base an agent is tempted to resolve a relative path against whether or not that base is the right one [10]. Occasionally the same disease showed a different symptom, an EISDIR error, the agent opening a directory as if it were a file [14].
This is why a debugger was the wrong instrument. Mixed path conventions in ordinary deterministic software fail consistently, with the same stack trace every run [16]. Here the same phase on the same kind of input passed on Tuesday and failed on Wednesday: identical code, identical input file, different outcome per run [11][17]. Searching the symptom returns the classical diagnoses, race conditions, temporary files deleted too early, a directory that did not exist yet, and none of them describe an executor that reads the same contract twice and resolves it differently [15].
Two mechanisms kept the whole thing off the books. Retries absorbed most of it: a failed read was retried, the agent tried another path, found the file and moved on, and articles kept arriving in the morning, which the author calls retry masking converting failures into costs [12][13]. And the run reports written after every run captured this error exactly zero times, because the failures lived one level down in the session transcripts, where a retried error leaves a trace but no alarm [18]. That is nought percent monitoring coverage over 117 known occurrences [21]. At hundreds of agent sessions a week, the author's earlier work on prompt caching costs is the relevant frame: per-call failures add up to real money and real hours [8][19].
The detection story is the useful part. What cracked it open was not a debugger but counting [5]. Averaged over the 49-day window, the error fired about 2.4 times a day, roughly 17 times a week [20][22], a rate low enough to look like weather and high enough to pay for.
Watch whether transcript-level error classes get promoted into run reports at all, since anything a retry survives is currently invisible in production [18]. Watch retry counts per phase as a first-class metric rather than a debugging afterthought [12]. And watch how briefs express paths, because the failure is in the contract between code and interpreter, not in either side's source [23].