Skip to content

Build1 publisher3 min readPublished

Strands and CrewAI agents reported success on runs where verification never executed

Strands and CrewAI agents exited 0 on all six runs where a broken verification tool checked nothing, a recorded-proxy test on dev.to found. Only tool-call traffic captured outside the framework showed the failures.

The Engineer · Build desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

Illustration accompanying Strands and CrewAI agents reported success on runs where verification never executed
Generated illustration

What happened

  • LangGraph, whose tool calls live in code, stopped with a hard error on all three runs of the same broken-verifier test.
  • Strands runs where the verification tool answered with processing errors such as "document 88 not found" still exited 0 every time.
  • When a required argument was added without telling the model, Strands and CrewAI accepted six invented values, and all three CrewAI runs invented the same sentence.
  • Two Strands runs finished with reason stop and an empty answer while the 100- and 106-word drafts sat in the last tool-call argument.
  • In a separate experiment, an irreversible publish ran twice under two different call IDs, and both calls returned success.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • decision A CI gate for agent runs has to count executed verification calls in recorded traffic, because in this test the exit code only reported that the loop had ended.
  • constraint Deduplicating on call ID or argument bytes cannot stop a reworded retry from repeating an irreversible action, so the idempotency key has to come from the task, which the model cannot rewrite.
  • cost Silent failures keep spending model calls on every run, and nobody reviews the spend because the run shows green.

"The failures that cost me time were never the crashes: a crash is a gift, because CI sees it," the post's author wrote on dev.to [19]. The test ran one small task (fetch headlines, write a draft, verify it through a tool) in Strands, CrewAI and LangGraph, all behind a single recording proxy that kept every attempt [1].

Which framework failed loudly depended on where the tool calls get decided. LangGraph's tool calls live in code [5]. In the other two, the author broke the verifier so it returned a static count instead of checking the draft. The tool still ran and answered, and the run reported success [2]. "LangGraph is the only one that cannot slip through silently: its structure cannot step over a broken tool," the author wrote [6]. The empty-answer failure turned up only in the model-driven frameworks [11]. "A framework whose output node is the deliverable cannot produce this shape," the post says [12].

Tool errors did not reach the exit code either. Strands logged eight tool error responses across its three runs of that test, and every process finished green [2]. "The error responses existed on the wire, and the tool layer even counted them. The process just never looked," the author wrote [13].

The invented-argument result comes down to one line of validation. The new required argument was checked only for being a non-empty string [9]. A validator like that confirms the model typed something [9]. A plausible sentence about word counts cleared it with zero errors in both model-driven frameworks [8].

A trace only helps if it is recorded at the tool boundary. The framework's own trace did not show the duplicate calls. Each one carried a different call ID, so on the wire they looked like two unrelated requests [17]. The double publish came back as two successes for that reason [15]. An idempotency ledger keyed on the argument bytes misses a retry that rewords the same intent [18].

The failures also cost money. The runs that did nothing still spent 15 LLM calls in CrewAI and 13 in Strands, summed over three runs each [3]. On crash recovery, the framework without state recording re-ran the whole task for 2 LLM calls. The one with a checkpointer resumed with 0 [14].

Now, how far does this transfer. Three runs per framework on one task is enough to show that exit 0 can sit alongside zero executed verifications [1]. It is not enough to rank Strands against CrewAI. The author wrote that the contrast on invented arguments is "not purely a framework property", because the two frameworks advertised the new parameter differently on the wire and did not sample the same way [10]. I'd expect the exit-code finding to hold in any setup where the model decides whether to call the verifier and the process reports only that the loop ended. The code, traces and analysis scripts are on GitHub [20].

What to watch

  • A rerun of the published harness with more than three runs per framework and the new parameter advertised identically on the wire.
  • Whether Strands or CrewAI start surfacing tool error responses in the process exit status or run result.
  • Whether a dedupe key derived from the task stops the double publish in the author's side-effect experiment.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories