Build1 publisher3 min readPublished
The retry loop is a memory bug: a gatekeeper that checks failures before it calls the tool
A runnable Mem0-backed wrapper blocks an agent's repeat tool call before it burns the request. The pattern is sound; the default policy will block you forever.
The Engineer · Build desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction
What happened
- An agent calls a flaky API, gets a 429, retries with the same arguments, and is rate-limited again; ten minutes later in a fresh session it repeats the same call, because the lesson lived in a transcript that was thrown away when the process exited.
- A post about manually gatekeeping AI agent tool calls hit the front page of dev.to this week with 48 comments, mostly developers arguing over how much an agent's tool-calling loop can be trusted.
- The article presents a MemoryGatekeeper class that sits between an agent's decision to call a tool and the execution: before calling, search memory for semantically similar past attempts; if a similar attempt failed recently, block the call and return the remembered reason instead of burning a real API call; after every call, success or failure, write the outcome back to memory.
- Requirements are Python 3.10+, an OpenAI API key (Mem0's default extraction pipeline uses it to turn raw text into structured memories), and the mem0ai package installed via pip; no separate Mem0 account is required for this local setup.
- Mem0 defaults to a local, on-disk vector store, so the example does not talk to a hosted Mem0 service and is a self-contained memory layer.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
A dev.to walkthrough published this week ships a small, complete `MemoryGatekeeper` class that sits between an agent's decision to call a tool and the actual execution: search memory for similar past attempts first, block if a similar attempt failed, write the outcome back either way [3]. The failure mode it targets is a billing line, not a prompting problem: an agent hits a 429, retries with identical arguments, gets rate-limited again, and then repeats the whole sequence in a fresh session because the lesson lived in a transcript that was discarded when the process exited [1].
The framing is worth noting because the same week a post about manually gatekeeping agent tool calls reached the dev.to front page with 48 comments, mostly developers arguing about how far an agent's tool-calling loop can be trusted [2]. This is the version of that argument with code attached.
The mechanics are unremarkable, which is the point. `check()` builds a query string of the form "tool call: {name} with args {args}" [14], runs `memory.search` with `limit=3` against the agent's memory namespace, and returns an (allowed, reason) pair [7]. It blocks when a hit scores at or above a `block_threshold` that defaults to 0.75 and the remembered text contains the substring "failed" [6]. `record()` appends a line stating whether the call succeeded or failed, a UTC ISO timestamp, and a detail string [8]. The wrapper prints a `[BLOCKED]` line and returns `{"blocked": True, "reason": reason}` instead of calling the function, and on a real exception it records the failure and re-raises [9].
Setup is Python 3.10+, the `mem0ai` package, and an OpenAI key, which Mem0's default extraction pipeline uses to turn raw text into structured memories; no separate Mem0 account is needed [4]. Mem0 defaults to a local on-disk vector store, so the memory layer does not call a hosted service [5]. The demo is a fake API that raises `RateLimitError("429: rate limit exceeded, retry after 3600s")` for the endpoint `/reports/daily` [11]. First run fails and records; second run of the same script prints the block with the remembered failure text and timestamp [12]. No second API call, no second rate-limit hit [13].
Two things to understand before you put this in a hot path. First, the block is semantic, not exact: the post notes that `fetch_weather(city="NYC")` and `fetch_weather(city="New York")` will match each other, which an exact-match cache would miss [10]. That is the real argument for the vector store over a dict, and it is also the reason a 0.75 threshold is a policy decision you now own rather than a constant you copy.
Second, the code remembers the retry-after but never reads it. `check()` compares score and substring only; nothing compares the recorded timestamp against now [6][7][8], so a single recorded failure blocks that call shape indefinitely, long past the 3600 seconds the error message asked for [1]. A transient 429 becomes a permanent refusal until you delete the memory. Add a TTL or parse the retry-after before shipping.
On cost, an allowed call carries one search plus one add; a blocked call carries the search alone [2]. With Mem0's extraction pipeline in the path, "cheap lookup" is relative to the API call you avoided, not free [4].
Watch whether the block policy grows a recency term, and whether the substring test survives contact with real error strings: a successful call records "succeeded" with detail "ok" and will not trip the filter, but any detail text containing the word "failed" will [3].