Build1 publisher3 min readPublished
Agentic planning added 3 of 32 exact-match points in a 100-question GraphRAG benchmark
A TigerGraph hackathon entry raised exact match from 67% to 99% on 100 questions, with the agent alone accounting for 3 of the 32 points. The rest needed a parser that made Wikipedia infobox fields countable in the graph.
The Engineer · Build desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened
- Plain RAG and GraphRAG with entity linking and one-hop traversal tied at 67% exact match, two of six pipelines run on the same model and the same questions.
- On the 21 aggregation questions, plain RAG answered one correctly and GraphRAG answered none.
- A second graph layer parsed from the articles' infoboxes, built with zero LLM calls, took aggregation from 0 of 21 to 21 of 21.
- Replacing the generative planner with two typed selection calls from TypeSafe's System One held 99% and cut time per question from 13 seconds to 1.7.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- decision Teams building agentic GraphRAG get more from modelling countable, filterable fields first; on this benchmark, agent planning paid off only once those fields existed in the graph.
- cost On questions the graph already models, a generative planner buys no accuracy over typed selection, so its latency and per-question model calls are overhead someone pays for.
- exposure Agent comparisons graded only by an LLM judge can credit fluent refusals as correct, so judge-only scores overstate how often a pipeline actually answers.
One aggregation question had 43 matching documents. Retrieval fetched the top five, and the model reported a count equal to the size of what it was shown [8]. The other 38 documents never reached the context [5], so no prompt change could fix the answer. An agent planning over the same retrieval gets more attempts at a search that was never built to return every match. I'd expect that to move the score a little. It did: the agent alone took exact match to about 70% [1].
The first graph did not help. It was LLM-extracted, with 12,000 entities and free-text relationship labels, and "competitors: 43" was never a property anyone could filter or count on [9]. "I had modelled the prose and not the facts," the author wrote [10]. The fix was a parser, and it is the best engineering in the writeup. Every article opens with an infobox listing event, games, venue, date, competitors, nations and gold medallist [11]. The author turned those fields into OlympicEvent vertices linked to Games, Sport and Venue, with a PREV_GAMES edge so "the Olympics before 2016" is one hop [11].
The headline split needs a qualification. "The agentic architecture alone bought 3 points. The agent with the right retrieval surface bought 32," the author wrote [14]. Those figures describe an interaction between two parts. The published text does not include a score for the infobox layer with no agent over it, so the 29 points between them [6] cannot all be credited to structure. The author argues the agent handles whatever the layer does not model. Asked who directed a film, it switches to vector search on its own [15].
Next the author asked whether the planner needs to reason at all. The test started from the premise that the agent picks the structured query on 100 of 100 questions [16]. Every decision the planner makes already exists as a row in the graph: 47 sports, 20 Games, 316 venues and 475 event names [16]. "So planning isn't generation, it's selection," the author wrote [17].
The typed path is about 7.6 times faster [2], but its call count matters more. The author's free tier allowed 500 calls a day, enough for one benchmark run, and measurements hit that limit three times in one day [19]. A path with no generation calls cannot hit a rate limit halfway through a live demo [19].
The question mix decides whether any of this transfers. Aggregation is 21 of the 100 questions [4], and the 21-answer swing on that type is about two-thirds of the 32-point spread [3]. The corpus is 2,951 Wikipedia articles [3] built around uniform infoboxes [11]. For the result to hold elsewhere, the documents need a field block a parser can read, and users need to ask for counts and filters over it. Every pipeline used the same cheap model, gemini-3.1-flash-lite [4]. A larger model would still see only the documents retrieval hands it [8].
Few hackathon entries report where their own grader failed. This one does. The LLM judge gave 4 or 5 out of 5 to 14 wrong answers, mostly fluent refusals like "the corpus does not contain this" [21]. Exact match, computed with no model in the loop, is the headline metric for that reason [5][21].
The author concludes that agents earn their cost on open-ended questions and are overkill where a typed query settles it. The agent still ships, because nobody knows in advance which kind of question is coming [20]. I think keeping it is right for a corpus where the planner's entire decision space comes to 858 rows [7].
What to watch
- The repo's full six-pipeline table: a score for the infobox layer without agent planning would show how much of the 29 points belongs to structure alone.
- Whether the typed-selection path holds 99% on a question set weighted toward open-ended questions outside the modelled Olympic fields.