BuildNot yet confirmed elsewhere1 publisher2 min readPublished
Postman found tool-selection errors rising past about 40 tools visible to its agent
Postman says Agent Mode, its AI agent for 40 million developers, made more tool-selection errors once the model could see more than about 40 tools. The team now limits the model to the tools each task needs, and found the harder work in APIs built around its interface.
The Engineer · Build desk

What happened
- Postman's team expected model quality and prompt design to be the hardest problems and found the deeper ones in fitting an agent into a mature product.
- Early atomic tools made long workflows feel slow, because every tool call had to return to the model before the next step could begin.
- Larger or newer models reduced the tool-selection errors in Postman's testing but did not eliminate them.
- Many Postman client APIs were tied to interface state: some tools needed certain elements open, and others opened new tabs as side effects.
- Agent Mode asks for user approval before any action that modifies application state.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- constraint In Postman's testing a flat tool list stopped scaling at a few dozen entries, so an agent over a large product needs a per-task selection layer before it needs more tools.
- decision Products whose APIs assume an open tab or panel must choose between reworking those APIs for direct data access and letting the agent drive the interface, side effects included.
- exposure Zero data retention for Agent Mode on Bedrock is model-dependent, so a model swap made for cost or quality can change whether customer data is retained.
Postman started with atomic tools for a defensible reason. Small actions, such as opening a request or fetching one piece of metadata, gave the team correctness and control in early iterations [10]. Fine-grained tools also multiply, and the post ties selection errors to how many tools the model can see at once [12].
I would treat Postman's figure of roughly 40 as a property of its own catalogue. One failure the post describes is the agent choosing tools that "seemed semantically reasonable but were wrong in context" [13]. How often that happens depends on how much the tools overlap in name and purpose. Forty distinct tools and forty near-duplicates are different selection problems. The threshold carries over to another product only if that product's tools overlap about as much and its agent runs on comparable models.
The replacement architecture picks tools by need and context and isolates individual execution threads, so the model sees only what the current task requires [15]. The post files this under treating context as the primary bottleneck. That is one of three patterns it names, alongside controlling tool sprawl and exposing schema-based reads [3].
The coupling to interface state has a history. Postman has evolved over 11 years, and its users learned to find information by expanding sidebars, checking tabs and opening requests [5]. The post puts the agent's position plainly: "An agent reasons over data rather than navigating a screen." [6] An API that expects an open tab therefore makes the agent operate the interface before it can change anything [16]. I would expect schema-based reads [3] to be Postman's answer to this. The available text ends before that section, and it does not name the models tested or report error rates.
The post is candid about where its controls stop. Postman scopes tools to the task and selects purpose-built context, and says these controls reduce unintended actions and data exposure while production testing and monitoring remain necessary [17]. Redaction of personally identifiable information through Amazon Bedrock Guardrails runs before text reaches the model, and enterprise admins turn it on in Agent Mode's guardrail settings [9].
AWS co-wrote the post [3], and its Bedrock section is about operations. Postman describes variable, latency-sensitive demand with sharp traffic bursts. It says Bedrock lets it scale that workload without operating its own model-serving infrastructure [7]. The features it names include geographically scoped cross-Region inference and multi-tier prompt caching [4].
What to watch
- Error rates by toolset size and model from Postman would show whether the 40-tool threshold comes from the models or from overlap in its catalogue.
- Whether Postman's schema-based reads rework its interface-bound client APIs or wrap them for the agent.
- Whether the 40-tool threshold moves as Postman changes models through Bedrock.
Clarity's read
What the record supports and how the coverage leans. The claims behind it follow.
Reality
- Evidence40
- Adoption35
- Hype gap+15
- Incentives80
- Confidence50
Claim ledger
Ranked by verification strength, evidence, and original report placement.
- [1]
Postman runs Agent Mode, an AI-native way to work across API testing, documentation, discovery and implementation, for 40 million developers on Amazon Bedrock.
- [2]
The Postman team expected model quality and prompt design to be the hardest problems; the deeper challenges came from integrating an agent into a mature product with years of interface-driven assumptions, a wide surface area and specialized concepts.
- [3]
In the post, Postman and AWS describe architectural patterns including controlling tool sprawl, exposing schema-based reads, and treating context rather than capability as the primary bottleneck.
- [4]
Agent Mode uses Amazon Bedrock for model flexibility, geographically scoped cross-Region inference, model-dependent zero data retention, and multi-tier prompt caching.
- [5]
Postman has evolved over 11 years, and developers and users learned to locate information through the interface by expanding sidebars, checking tabs and opening requests.
- [7]
Postman's global developer community creates variable, latency-sensitive demand with sharp traffic bursts; with Amazon Bedrock, Postman can scale the workload without operating its own model-serving infrastructure.
- [8]
Agent Mode requires user approval before actions that modify application state.
- [9]
Postman uses Amazon Bedrock Guardrails to redact personally identifiable information before it reaches the underlying LLM; enterprise admins can turn this on in Agent Mode's guardrail settings.
- [10]
Early on, the team leaned toward highly atomic tools, such as opening a request, updating one field or fetching a specific piece of metadata, an approach that supported correctness and control in early iterations.
- [11]
Many real-world workflows required long sequences of tool calls; the experience felt slow because every action had to return to the model before the next could begin, and users watched the agent step through actions they had mentally grouped as a single operation.
- [12]
In Postman's testing, tool-selection errors increased once the visible toolset exceeded approximately 40 tools.
- [13]
The agent could call nonexistent tools, pass incorrect arguments despite valid schemas, or select tools that seemed semantically reasonable but were wrong in context.
- [14]
Larger or newer models reduced the tool-selection errors but did not remove them.
- [15]
The current architecture selects tools based on need and context and isolates individual execution threads; the model sees only the tools relevant to the current task.
- [16]
Many client APIs were implicitly coupled to interface state: tools that modified requests needed certain elements to be open, while other tools opened new tabs as side effects, and the agent had to open a request tab.
- [17]
Postman scopes available tools to the task, selects purpose-built context and applies model-dependent data-retention settings; it says these controls reduce unintended actions and unnecessary data exposure, while production testing and monitoring remain necessary.
Sources
1 independent publisher whose own reporting we read for this story.
- aws.amazon.comHow Postman runs Agent Mode for 40 million developers on Amazon Bedrock
1 article · October 9, 2026
Topics and entities
Follow any of these and your For You feed starts watching them — no settings page required.
Topics
- AI AgentsFollow
- LLM context managementFollow
- LLM tool and function callingFollow