Skip to content

Build1 publisherNot yet confirmed elsewhere2 min readPublished

Search agents keep calling tools after judging the results useless, a 90,000-episode study finds

Researchers at NTU, NUS and CMU found seven models judged useless search results correctly 97% to 100% of the time yet kept searching anyway. For operators, a spending cap holds only when the runtime enforces it by taking the tools away.

The Engineer · Build desk

How we use AISend a correction

Illustration accompanying Search agents keep calling tools after judging the results useless, a 90,000-episode study finds
Generated illustration

What happened

  • The pre-registered study ran 90,000 episodes in a controlled retrieval-failure environment built to separate how agents judge evidence from how they decide to stop.
  • Adding an explicit financial cost per tool call to the prompt shifted stopping time slightly but did not link stopping to the quality of the evidence.
  • Stating a step budget in the prompt moved the stopping distribution of 7B to 8B open models onto the final deadline.
  • Telling models they could answer from internal memory caused arbitrary early stops, whether or not external retrieval had worked.
  • A harness that removed tool access after five consecutive useless results, leaving only the answer action, raised task success for every model and held its stopping point when the budget doubled.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • exposure Agents deployed the way the post describes enterprise setups, with a budget and cost warnings in the system prompt, rely on a control that the models' stopping behaviour did not follow in the study.
  • decision Whether a small open model should see its step budget at all is now a design question, since seeing the number gave those models a target to spend toward.
  • capability A stop enforced in harness code lets an operator raise the step ceiling for hard tasks without paying for longer dead-end searches.
  • cost Every turn an agent spends past a dead end is billed in tokens, API fees, latency and context degradation, and the operator pays it whatever the prompt says.

The dev.to post that summarises the study offers its own explanation. An autoregressive model "has no endogenous balance sheet" and "experiences zero disutility from spending compute," the post argues [8]. A final answer can be scored as wrong, and one more search defers that moment, a move the post calls "risk postponement" [11]. The model, unlike the operator, never sees the invoice [8]. This is the post writer's interpretation of the results.

The measurement deserves more attention than the theory. The researchers built a time-matched contrast. It asks whether an agent stops more often after an unbroken run of useless results than after a run that contained useful evidence [3]. For agents governed by their prompts, the contrast stayed near zero [3]. The same models were near-perfect at labelling those results as useless [2]. The judgment is present in the output, and the stopping decision ignores it. Separating the two is careful experimental design. A cruder eval would lump an agent that cannot tell bad results from good together with an agent that can tell and keeps searching anyway.

The harness rule that fixed this is small [7]. The code is a counter of consecutive useless results and a branch that drops every tool except the answer action. The write-up does not say how the harness decided a result was useless, which seven models were tested, or how large the success gain was. In the study, the environment controlled which retrievals failed [1]. A production harness has to get that label from somewhere.

I'd take it from the model, because the model is accurate at this job [2]. Have it emit a usefulness grade for each tool result as a structured field. Let the harness do the counting and the cutting. For retrieval work, where a result either answers the question or misses, I think that split is right: the model judges, and code decides when the judging has gone on long enough [7].

For the numbers to transfer, production failures have to resemble the study's. The rule trips only on an unbroken run [7]. A search tool that returns an occasional near-miss resets the counter each time it does, and the agent keeps spending up to whatever ceiling the runtime sets. The evidence also comes from one controlled retrieval-failure environment and seven model architectures, so anyone running a model outside that set is extrapolating [1].

What to watch

  • The full paper's per-model results, including which seven architectures were tested and how much the five-result rule raised task success.
  • A replication on production retrieval traffic, where near-miss results could keep a consecutive-failure counter from ever reaching five.
  • Whether agent runtimes ship consecutive-failure stop rules as a default setting.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories