Skip to content

Build1 publisher3 min readPublished

Three workflow bugs made one team's GPT-5.4 agent look lazy in production

One team running a GPT-5.4 agent in n8n traced its production 'laziness' to three workflow bugs, the first a retry cap cut from 6 to 2. Fixing the loop restored quality on the same model, so traces and stop reasons should be checked before any model swap.

The Engineer · Build desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

Illustration accompanying Three workflow bugs made one team's GPT-5.4 agent look lazy in production
Generated illustration

What happened

  • An n8n branch marked any run successful if it returned valid JSON with an answer longer than 280 characters.
  • The team's OpenAI-compatible API path accepted the first acceptable-looking response instead of the best one.
  • Staging traces showed retrieval, comparison and verification steps, while production traces showed one weak source followed straight away by an answer.
  • The shallow answers still looked polished, and the post says they passed casual review more often than they should have.
  • Run through the broken production scaffold, GPT-5.4 looked bad, and under the staging flow the same model looked fine.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • constraint One global retry cap sized for classification will run research and debugging agents out of budget, so caps have to be set per task type against the length of each tool chain.
  • exposure A team that scores only final outputs can ship confident wrong answers while its success gate keeps reporting green runs.
  • decision A decision to blame or replace a model holds up only after the same scaffold has been rerun with one variable changed. A production-versus-notebook comparison changes too much to isolate the model.

The three faults stack. The retry cap fell by two-thirds [1]. The post argues that two attempts is often a trap for research, debugging or any tool use. One bad retrieval plus one tool hiccup leaves the agent out of budget [10]. Its example of the drift is a config small enough to miss in review [11]:

```json { "task_type": "research", "max_retries": 2, "timeout_seconds": 20 } ```

A run that survives the budget then reaches the success gate. The post gives it as pseudo-logic [12]:

```js const passed = isValidJson(response) && response.answer.length > 280; ```

That line never asks whether a search ran or a claim was checked. "That is a shallow-answer reward function," the author wrote [13]. The third fault took away any second look at the output [6]. With all three in place, the cheapest output that parses and clears 280 characters is the one that ships.

The author's phrasing is "This is how you accidentally train an agent to stop early." [22] The word is loose. The model family was the same in staging and production [8]. What differed was which outputs the workflow counted as done. A vendor swap would leave that gate in place. The post says the same failure can show up whether the router points at GPT-5.4, Claude Opus 4.6 or Grok 4.20 [20].

The diagnostic work is good engineering. The team kept the broken production scaffold fixed and ran different model families through it [17]. That is the controlled test the post recommends: identical system prompt, tool definitions, retry budget, output limits, evaluator, API path and success criteria, with one variable changed at a time [14]. Comparing production n8n against a clean notebook script, by the post's account, changes the entire experiment [14].

Stop reasons told the team more than final-answer scoring did [15]. For Anthropic agents the post lists end_turn, max_tokens, tool_use and pause_turn [15]. For OpenAI-compatible stacks it says to check whether the expected tool calls fired [16]. The question for each run is whether the model finished or the orchestration layer decided it had enough [16].

The evidence is one team's incident. The post does not report run counts, error rates or scores before and after the fix. It says only that quality came back [7]. It then generalises: "Most of the time, the model did not suddenly get worse." [19] That claim rests on one workflow. For it to hold elsewhere, a stack needs at least one of the same conditions: a retry budget shorter than its tool chain, a success check on format and length, or acceptance of the first response. The post called its root cause "boring and very fixable" [21], and incident reviews rarely get a kinder sentence. In my view those three conditions are what to look for in the traces before filing a regression against a model. In the author's words, "A strong model inside a bad loop will look worse than a decent model inside a clean loop." [18]

What to watch

  • Whether Claude Opus 4.6 and the other model families the team tested degrade under the same production scaffold the way GPT-5.4 did.
  • Before-and-after figures from the team: tool calls per run, stop-reason mix and error rates once the retry cap and success gate were fixed.
  • Whether agent orchestration tools such as n8n start treating expected tool calls and stop reasons as part of the default success check.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories