Skip to content

BuildNot yet confirmed elsewhere1 publisher2 min readPublished

Postman found tool-selection errors rising past about 40 tools visible to its agent

Postman says Agent Mode, its AI agent for 40 million developers, made more tool-selection errors once the model could see more than about 40 tools. The team now limits the model to the tools each task needs, and found the harder work in APIs built around its interface.

The Engineer · Build desk

How we use AISend a correction

Illustration accompanying Postman found tool-selection errors rising past about 40 tools visible to its agent
Generated illustration

What happened

  • Postman's team expected model quality and prompt design to be the hardest problems and found the deeper ones in fitting an agent into a mature product.
  • Early atomic tools made long workflows feel slow, because every tool call had to return to the model before the next step could begin.
  • Larger or newer models reduced the tool-selection errors in Postman's testing but did not eliminate them.
  • Many Postman client APIs were tied to interface state: some tools needed certain elements open, and others opened new tabs as side effects.
  • Agent Mode asks for user approval before any action that modifies application state.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • constraint In Postman's testing a flat tool list stopped scaling at a few dozen entries, so an agent over a large product needs a per-task selection layer before it needs more tools.
  • decision Products whose APIs assume an open tab or panel must choose between reworking those APIs for direct data access and letting the agent drive the interface, side effects included.
  • exposure Zero data retention for Agent Mode on Bedrock is model-dependent, so a model swap made for cost or quality can change whether customer data is retained.

Postman started with atomic tools for a defensible reason. Small actions, such as opening a request or fetching one piece of metadata, gave the team correctness and control in early iterations [10]. Fine-grained tools also multiply, and the post ties selection errors to how many tools the model can see at once [12].

I would treat Postman's figure of roughly 40 as a property of its own catalogue. One failure the post describes is the agent choosing tools that "seemed semantically reasonable but were wrong in context" [13]. How often that happens depends on how much the tools overlap in name and purpose. Forty distinct tools and forty near-duplicates are different selection problems. The threshold carries over to another product only if that product's tools overlap about as much and its agent runs on comparable models.

The replacement architecture picks tools by need and context and isolates individual execution threads, so the model sees only what the current task requires [15]. The post files this under treating context as the primary bottleneck. That is one of three patterns it names, alongside controlling tool sprawl and exposing schema-based reads [3].

The coupling to interface state has a history. Postman has evolved over 11 years, and its users learned to find information by expanding sidebars, checking tabs and opening requests [5]. The post puts the agent's position plainly: "An agent reasons over data rather than navigating a screen." [6] An API that expects an open tab therefore makes the agent operate the interface before it can change anything [16]. I would expect schema-based reads [3] to be Postman's answer to this. The available text ends before that section, and it does not name the models tested or report error rates.

The post is candid about where its controls stop. Postman scopes tools to the task and selects purpose-built context, and says these controls reduce unintended actions and data exposure while production testing and monitoring remain necessary [17]. Redaction of personally identifiable information through Amazon Bedrock Guardrails runs before text reaches the model, and enterprise admins turn it on in Agent Mode's guardrail settings [9].

AWS co-wrote the post [3], and its Bedrock section is about operations. Postman describes variable, latency-sensitive demand with sharp traffic bursts. It says Bedrock lets it scale that workload without operating its own model-serving infrastructure [7]. The features it names include geographically scoped cross-Region inference and multi-tier prompt caching [4].

What to watch

  • Error rates by toolset size and model from Postman would show whether the 40-tool threshold comes from the models or from overlap in its catalogue.
  • Whether Postman's schema-based reads rework its interface-bound client APIs or wrap them for the agent.
  • Whether the 40-tool threshold moves as Postman changes models through Bedrock.

Clarity's read

What the record supports and how the coverage leans. The claims behind it follow.

Reality

Evidence40
Adoption35
Hype gap+15
Incentives80
Confidence50
Why these scores

Claim ledger

Ranked by verification strength, evidence, and original report placement.

  1. [1]

    Postman runs Agent Mode, an AI-native way to work across API testing, documentation, discovery and implementation, for 40 million developers on Amazon Bedrock.

    ReportedSupportedView cited source
  2. [2]

    The Postman team expected model quality and prompt design to be the hardest problems; the deeper challenges came from integrating an agent into a mature product with years of interface-driven assumptions, a wide surface area and specialized concepts.

    ReportedSupportedView cited source
  3. [3]

    In the post, Postman and AWS describe architectural patterns including controlling tool sprawl, exposing schema-based reads, and treating context rather than capability as the primary bottleneck.

    ReportedSupportedView cited source

Sources

1 independent publisher whose own reporting we read for this story.

  1. aws.amazon.com

    1 article · October 9, 2026

    How Postman runs Agent Mode for 40 million developers on Amazon Bedrock

Share your take

Let Clarity write the post for you.

Signed-in readers get a short post drafted on this story in the register they choose — narrative, analytical, or a direct position — editable to the last word before it goes anywhere. The share buttons at the top of this story work without an account.

Topics and entities

Follow any of these and your For You feed starts watching them — no settings page required.

Loading related stories