Build1 publisher3 min readPublished
Rate limit your MCP servers, because a retrying agent turns one error into a billing incident
A dev.to walkthrough on agentgateway makes a narrow point worth taking seriously: an agent that retries on failure will keep hitting an unrated MCP server until something gives.
The Engineer · Build desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened
- Without controls on how many times an agent or LLM can hit an MCP server, you are exposed to potential DOS attacks, memory hogging, insane API bills, and server/system overload.
- LLMs retry if an error occurs; the author says this is part of their programming so that the user gets at least some answer, even if it is not correct.
- A non-rate-limited loop of MCP tool calls can cause huge API bills or crash the system, because the agent hitting the LLM/MCP server will eat up all available memory.
- The author's analogy is an application memory leak: a bug that lets software keep consuming RAM it no longer needs until the system runs out of memory and a crash occurs.
- When a client or agent begins a request it hits tools/list one time to learn which MCP tools are available.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
A walkthrough published on dev.to argues that an MCP server with no cap on how many times an agent or LLM can call it is exposed to denial of service, memory hogging, large API bills and general server overload [1]. The mechanism the author names is the part operators should care about: LLMs retry when a call errors, because returning some answer is treated as better than returning none, so one failing tool call can become a loop nobody scheduled [2][3].
The author's analogy is a memory leak: software that keeps taking resources it no longer needs until the host runs out and something crashes [4]. Applied to MCP, the claim is that an unlimited retry loop either runs up the API bill or consumes available memory on the machine serving the agent [3].
The traffic shape matters for where that load lands. According to the post, an MCP client calls tools/list once, caches the tool schemas locally by name, description and input parameters, and thereafter picks tools from that cache rather than re-listing, which would waste tokens [5][6][7]. A tool call itself is a plain HTTP POST to /mcp carrying a JSON-RPC tools/call, answered with a 200 and a result [8]. If discovery happens once per session and is cached, then retry volume does not spread across the protocol; it concentrates on tools/call [9].
There is also a control-plane gap the post identifies. LLM rate limiting can be expressed as requests per window or as tokens per window, for example 100 tokens a minute, while MCP rate limiting is about the number of requests to the MCP server in a given window [10][11]. A token budget on the model therefore does not bound how many times the MCP server is hit, because the two limits count different units [12].
The implementation shown uses agentgateway on a Kubernetes cluster, with Kind or Minikube named as sufficient, plus a GitHub account [13]. It is two pieces of configuration: a gateway that fronts the MCP server, and the rate limiting policy itself [14]. The example targets the GitHub Copilot MCP server on the grounds that most readers already have GitHub access [15]. The plumbing is an Opaque Secret holding a GitHub PAT as an Authorization bearer value in the agentgateway-system namespace [16], a Gateway with gatewayClassName agentgateway listening on port 3000 over HTTP and accepting routes from the same namespace [17], and an AgentgatewayBackend of apiVersion agentgateway.dev/v1alpha1 with stateless session routing to api.githubcopilot.com on port 443, path /mcp/, over StreamableHTTP, with TLS and the PAT secret attached [18].
Two honesty notes. The retry premise is asserted, not measured: the post names no specific client, no retry count, and no recommended requests-per-minute figure [19]. And the material available here stops at the routing step, before the rate limiting policy is shown, so the policy syntax is not something I can describe [20].
What to watch is what the limit does to the caller. If an agent retries on any error [2], a rate limit rejection is also an error, so a gateway cap bounds the requests you forward and pay for, not the requests attempted [21]. Instrument both counters separately, or you will read a flat downstream graph as a fixed loop. Watch the PAT handling too: the example exports the token into a shell variable and pipes it into a heredoc [22], which is fine for a demo and a credential-in-history problem in a terminal you share.