Skip to content

Build1 publisher2 min readPublished

KAIST and Microsoft cut a grep agent's input tokens 57% by mapping named entities first

KAIST and Microsoft researchers say a map of named entities cut a search agent's input tokens from 206,500 to 88,100 per query on EnterpriseRAG-Bench. Whether that saving holds outside the benchmark depends on what the map costs to build and keep current.

The Engineer · Build desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

Illustration accompanying KAIST and Microsoft cut a grep agent's input tokens 57% by mapping named entities first
Generated illustration

What happened

  • Offline, CorpusMap writes one page per recurring entity that sums up its facts and holds verified backlinks to every document that mentions it.
  • Answer correctness on EnterpriseRAG-Bench rose from 62.1% with raw grep to 73.8% with the map, with GPT-5.5 as the agent's model.
  • On WixQA, per-query input fell from 337,200 tokens to 74,500, a 78% drop, and accuracy rose from a 67.5% baseline.
  • A folder-index baseline pushed input to 423,000 tokens on EnterpriseRAG-Bench and past 1.1 million on WixQA, and correctness fell to 48.8%.
  • A free-form LLM-written wiki used 192,400 tokens on EnterpriseRAG-Bench but scored 56.7%, below plain grep over the raw files.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • cost The reported savings count per-query input only. The offline extraction pass is a separate bill, so the net gain depends on how many queries share one build of the map.
  • decision Anyone running a folder index or a free-form LLM wiki in front of a grep agent now has a benchmark case for removing it, since both scored below plain grep on EnterpriseRAG-Bench.
  • constraint The gain depends on questions that hinge on people, projects, systems or modules recurring across files. Corpora of standalone documents give the extractor little to link.
  • capability The authors say fixed GraphRAG-style retrieval cannot recover when its graph walk misses an edge. An agent that still greps raw files can search past a gap in the map.

Give an agent grep, find and cat over a flat directory and it starts each query with no pointer from one file to another [4]. A dev.to write-up of "Follow the Entities: A Corpus Map for Agentic Search" (arXiv:2609.37226) [1] says the model has to rebuild the relationships between documents from scratch, at inference time, on every query [4]. The write-up's own example is an agent that reached turn six before it learned that one project's approval slip, specifications and status report sat in three different folders. Getting there cost about 200,000 tokens [14].

Per-query counts understate the gap, because the cheaper configuration was also right more often. Divide input tokens by correctness on EnterpriseRAG-Bench and raw grep costs about 332,500 tokens per correct answer [1]. CorpusMap costs about 119,400 [2], roughly 64% less [3]. The free-form wiki spent fewer tokens per query than raw grep but about 339,300 per correct answer [4].

The folder-index baseline roughly doubled per-query input on EnterpriseRAG-Bench [5]. Reading twice as much to score lower takes some effort. The write-up's explanation is that folder trees follow team charts or file formats, so the agent reads redundant directory summaries before it reaches evidence [15]. The wiki failed differently. Without strict grounding against the raw files, free-form wikis suffer hallucinations and missing cross-references, according to the write-up [16].

The layer that won is the most constrained one. I think that is the right design. Pages exist only for entities that recur across independent documents: people, projects, systems, vendors, code modules [8]. Each backlink is verified against a source file [9]. The agent keeps its terminal tools and still opens the raw documents to check what a page claims [10]. The authors argue that fixed retrieve-then-generate pipelines such as GraphRAG and HippoRAG cannot inspect downstream sources once the initial graph walk misses an edge [7].

Whether the 57% transfers depends on the workload [11]. The numbers come from two benchmarks, EnterpriseRAG-Bench and WixQA [c2, c3], with GPT-5.5 as the model on the EnterpriseRAG-Bench run [11]. A similar saving elsewhere needs questions that turn on entities spread across several files. A corpus of standalone documents gives the extractor little to link. The map also has to stay current as files change, since a stale backlink sends the agent to the wrong document. And the build has to pay for itself. The write-up does not give the token cost of the offline extraction pass [9]. At 118,400 input tokens saved per EnterpriseRAG-Bench query [6], the break-even query count is that extraction bill divided by 118,400.

What to watch

  • A reported token cost for CorpusMap's offline extraction pass, the figure that sets how many queries it takes to break even.
  • Results on a corpus that changes daily, showing how fast entity pages and backlinks go stale and what incremental rebuilds cost.
  • A replication with a model other than GPT-5.5, to show whether the 57% cut depends on how that one model navigates.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories