Skip to content

Build2 publishers3 min readPublished

Mistral ships multi-step retrieval you can run on your own index, on your own hardware

Agentic Search gives a model five tools to keep looking instead of answering from the first batch of chunks. The deployment terms matter more than the benchmark chart.

The Engineer · Build desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

Illustration accompanying Mistral ships multi-step retrieval you can run on your own index, on your own hardware
Generated illustration

What happened

  • Mistral AI launched Agentic Search on August 20th, giving enterprise AI systems a multi-step retrieval loop for searching, opening and reading long documents instead of answering from the first batch of text chunks.
  • Agentic Search is available through Mistral Search Toolkit and through Libraries inside Studio and Vibe.
  • Mistral says customers can run the tooling in the cloud or on-premises and connect it to an existing search index.
  • On-premises deployment against an existing index is described as an important requirement for organizations that cannot ship contracts, filings or operational records to a third-party service.
  • Agentic Search lets the model decide what to inspect next through five tools: search, open, navigate, read and grep.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

Mistral AI launched Agentic Search on August 20th, a multi-step retrieval loop that lets a model search an index, open a document, navigate to a page, read it and grep inside it before answering [1][4]. The interesting part for regulated buyers is not the loop but the deployment terms: Mistral says the tooling runs in the cloud or on-premises and connects to an existing search index [3].

That is the constraint that has kept banks, insurers and government suppliers away from hosted retrieval. If the retrieval layer is a third-party service, the contracts, filings and operational records have to leave the building to be indexed [c3b]. Mistral frames its tooling as portable and open so customers can work inside their own isolation boundaries [17], which is consistent with the efficient-models, open-tooling, customer-control pitch the company has run since Arthur Mensch, Guillaume Lample and Timothee Lacroix founded it in 2023 [16]. Reusing an index you already built and already audited is a smaller procurement question than standing up a new one somewhere else.

The mechanics are unglamorous, which is a compliment. Five tools, no model-specific fine-tuning required according to Mistral [4][5], sitting on top of the Search Toolkit that entered public preview on May 28th and handles extraction, chunking, embeddings, indexing, retrieval and evaluation across PDFs, office files, spreadsheets, emails and plain text [6]. Mistral's worked example asks for total monthly U.S. national defense expenditures in 1953: one search returned bulletins covering only part of the year, and the loop searched again, found a February 1954 Treasury Bulletin with all twelve months, read the page and totalled $44,463 million [15].

The numbers are Mistral's own, not independently replicated [14]. On FinanceBench, run with default chunking and ranking [7], the company reports Z.ai's GLM-5.2 going from 26.7% one-shot to 86% with the full toolset [8], scored by an LLM judge on a 150-question public subset of a benchmark covering 368 SEC filings and about 53,900 pages [10]. Read the decomposition rather than the headline: 52.6 of the 59.3 points came from simply being allowed to search more than once, and all four navigation tools together added 6.7, about 11% of the gain [8][20][21]. Mistral Medium 3.5 shows the same shape, 47.3 points from iteration and 8.7 from navigation [9]. On OfficeQA Pro, 133 questions over 696 Treasury Bulletins and roughly 89,000 pages of scanned tables, GLM-5.2 went from 6.3% to 51.9% [13].

Where navigation earns its place is cost. Against a search-only agentic loop, token use fell 23.9% for Mistral Medium 3.5 and 33.7% for GLM-5.2 [11], and FinanceBench p90 latency dropped from 255 seconds to 154, a 39.6% reduction, with mean latency from 108 to 71 seconds [12][22][23]. Those absolute figures are the ones to plan around. A p90 of 154 seconds is a queued job with a status page, not a chat box, and it is the improved number [12].

What to watch: whether anyone outside Mistral reproduces the FinanceBench and OfficeQA Pro deltas on their own corpora, since both benchmarks are public documents rather than the messy internal repositories the pitch is aimed at [10][13][14]. Also watch what "connects to an existing index" costs in practice, because five tools acting on someone else's index implies permission checks per call that the accuracy chart does not price [3][4].

Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories