Skip to content

Build1 publisher3 min readPublished

Foundry IQ knowledge bases ship as MCP servers, and four behaviours break naive clients

Microsoft's agentic retrieval is now callable from any MCP client, so non-Microsoft agent stacks can stop rebuilding chunking, embedding and permissions. The edges are where the work moved.

The Engineer · Build desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened

  • Every Foundry IQ knowledge base is also a standalone MCP server exposing one tool, knowledge_base_retrieve.
  • Any MCP-compatible client can call the knowledge base tool, including LangGraph via langchain-mcp-adapters.
  • A knowledge base wraps one or more knowledge sources, and the agentic retrieval engine handles query planning, parallel execution, semantic reranking and optionally answer synthesis.
  • The tutorial describes the hand-rolled alternative as a chunker, an embedding job, a vector store, a retriever, a reranker and a permissions filter bolted on at the end, with every new agent getting its own copy and every copy drifting.
  • The hand-rolled retrieval pipeline the tutorial lists comprises six components per agent.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

A dev.to tutorial on grounding a LangGraph agent in Microsoft Foundry IQ makes a point that outlasts the integration: every Foundry IQ knowledge base is also a standalone MCP server, exposing exactly one tool, `knowledge_base_retrieve` [1]. Any MCP-compatible client can call it, including LangGraph via `langchain-mcp-adapters`, which means agent stacks built outside Microsoft's own can consume enterprise retrieval instead of reimplementing it [2].

What sits behind that endpoint is the part teams keep rewriting. A knowledge base wraps one or more knowledge sources, and the agentic retrieval engine handles query planning, parallel execution, semantic reranking and, optionally, answer synthesis [3]. The tutorial describes the alternative as a chunker, an embedding job, a vector store, a retriever, a reranker and a permissions filter bolted on at the end, with every new agent getting its own copy and every copy drifting [4] - six moving parts per agent [1]. The knowledge base is not scoped to one consumer: the same one can ground a LangGraph agent, a Foundry Agent Service agent and a Copilot integration simultaneously, which is why the author argues for naming them after topics such as `hr-policy-kb` rather than after whichever agent got there first [5].

The four places it does not behave like a normal retriever, per the tutorial: the MCP tool result has a different shape from the REST/SDK retrieve response [6]; bearer tokens expire, so a static headers dict fails about an hour into a long-running graph [7]; per-user permission filtering needs a second token, distinct from the service credential [8]; and the knowledge base is itself a planner, so you have two planners per turn and must decide who does what [9].

The identity split is the one with a blast radius. The application authenticates to the search service with a service identity, and the end user's identity travels optionally in a separate header so the engine can filter documents that user should not see; conflating the two is described as the most common cause of "why is everyone seeing everything" bugs [10]. Source type interacts with this: indexed sources such as Blob, OneLake or an existing search index are ingested, chunked and vectorized into an index on your search service, while federated sources such as remote SharePoint, the web and other MCP servers are queried live at retrieval time and never ingested [11].

The operational prerequisites are unglamorous and load-bearing. The querying identity needs the Search Index Data Reader role [12], managed identity support requires Basic tier or higher [13], and if the knowledge base specifies an LLM the search service needs a managed identity with Cognitive Services User on the Foundry resource [14].

Watch the API version, because it is a product decision disguised as a config string. Agentic retrieval is generally available in the 2026-04-01 REST API, while 2026-05-01-preview adds answer synthesis, configurable reasoning effort, the `messages` input, document-level permissions and sensitivity-label metadata [15][16]; those features need the prerelease `azure-search-documents` SDK [17]. Two traps follow. Both the Azure portal and the Microsoft Foundry portal expose preview-only behaviour regardless of what your code targets, so the portal is not a reliable rehearsal of production [18]. And the API version also changes MCP behaviour [19], which means a client that works against preview is not automatically a client that works against GA.

Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories