Build1 publisher3 min readPublished
Plain RAG and GraphRAG missed all 31 counting and superlative questions on a TigerGraph benchmark
Plain RAG and GraphRAG got none of 31 counting and superlative questions right on a 100-question TigerGraph hackathon benchmark, an entrant reports. A COUNT from the graph, set beside the evidence actually read, shows when an answer is incomplete.
The Engineer · Build desk

What happened
- Asked how many 2004 Olympic athletics events had more than 41 competitors, plain RAG answered 1 and GraphRAG answered 2, against a true answer of 20.
- Both pipelines answered in the same confident tone, and neither output showed how much of the matching data it had actually retrieved.
- The author's pipeline ships every answer with an Investigation Certificate, a JSON record that sets the graph's own COUNT beside the evidence the agent inspected.
- A router sorts each question into one of three completeness classes and picks the cheapest tool that can satisfy it.
- On the same 100 public questions, with the same live graph and Groq model, the author's pipeline answered 99 correctly.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- decision Teams that answer counting or superlative questions over structured data now have a measured case for routing them to a database COUNT and limiting the model to tie-breaks.
- exposure A plain RAG answer built from a tenth of the matching data looks the same to an analyst as a complete one, so a wrong count can reach a report with no flag on it.
- capability Completeness becomes something a user checks by comparing two integers in a JSON record, with no need to trust the model's tone.
"Similarity search returns what looks relevant, but a counting question needs everything that is," the author wrote, adding that "no value of k fixes that, because the right k depends on the answer you haven't computed yet" [6]. I'd change one word. The k this question needed was the size of the matching set, 43, and that size is itself a count [4]. Plain RAG saw 4 of those 43 events, or 9.3% [22]. GraphRAG's one-hop expansion reached 12, or 27.9% [3][1].
The graph already holds the count. TigerGraph can return how many events match `sport = athletics AND games = 2004 Summer` as a plain COUNT; the author calls that number the structural bound [8]. Aggregation and superlative questions go to the exhaustive class, which runs a GSQL structural scan returning the COUNT plus every match [23]. The certificate for this question, pub-045, sets `"structural_bound": 43` beside `"evidence_set_size": 43` and records `"stop_reason": "structural_bound_met"` [10]. It spent 0 tokens and 227 ms, because this question never calls the LLM [10]. The same check on plain RAG would show evidence 4 against a bound of 43, a visible fail [11].
The LLM is called only to break a tie or recover a missing field. It never counts [13]. Routing starts with 5 regex templates, with TypeSafe AI's Jev System One as fallback, a non-autoregressive model that classifies intent at 0 output tokens [14]. On 5 questions the templates did not cover, regex sent all 5 to unverified RAG and Jev routed all 5 correctly [14]. "That small test (5 queries) is all the evidence I have for the classifier, so I'd call it promising and unproven," the author wrote [15]. I agree with the label.
The single miss shows what the certificate covers. Question pub-060 was a date-and-venue tie among 37 candidate events at ExCeL; Jev picked one at 0.90 confidence and picked wrong, as the LLM had in an earlier run [18]. Its certificate reads `pass_with_llm_recovery`, not a clean pass [19]. The evidence was complete and the answer was wrong [19]. The author wrote that "a certificate vouches for the evidence, and answer correctness stays a separate claim, so a system should say which one it is making" [20]. The same limit applies before the tie-break. Bound and evidence come out of one structural scan [23]. I'd expect a wrong filter predicate to produce two numbers that agree with each other and both describe the wrong set.
The author's cost figures put the pipeline at 18 tokens against GraphRAG's 2,203, and 0.41 seconds against 15.5, at 2.3 times the accuracy [16]. "The cost was always in stuffing chunks into a context window, and the graph queries are nearly free," the author wrote [21]. For those numbers to transfer, a team's questions have to filter on fields the graph already stores, the way this one filtered on sport and games [8]. The benchmark came from a TigerGraph hackathon and ran against a live TigerGraph graph [1][16]. Facts that live in unstructured documents would first have to be extracted into typed fields, and the post does not measure that cost.
What to watch
- Whether Jev's routing holds up on more than the 5 un-templated questions it has been tested on so far.
- Whether the 99 of 100 result holds on benchmark questions beyond the 100 public ones.
- Whether the structural-bound check is applied to a corpus where the counted fields first have to be extracted from text.