Plain RAG and GraphRAG got none of 31 counting and superlative questions right on a 100-question TigerGraph hackathon benchmark, an entrant reports. A COUNT from the graph, set beside the evidence actually read, shows when an answer is incomplete.
Reality
- Evidence45
- Adoption
- Insufficient
- Hype gap+25
- Incentives55
- Confidence50
A TigerGraph hackathon entry raised exact match from 67% to 99% on 100 questions, with the agent alone accounting for 3 of the 32 points. The rest needed a parser that made Wikipedia infobox fields countable in the graph.
Reality
- Evidence45
- Adoption
- Insufficient
- Hype gap+15
- Incentives35
- Confidence45
Fraudagent, a TigerGraph challenge entry, settles 20 fraud alerts in code whose build fails on any LLM import. Its verdicts match bit for bit with no API key, so the calibration bug its authors blame for clearing fraud lives in code a test can reach.
Reality
- Evidence45
- Adoption5
- Hype gap+10
- Incentives55
- Confidence45
Sentinel's builder trained its alert model on fraud versus cleared cases, reaching 0.9465 AUC after three passes that over-called fraud on every case. The finished agent lets an LLM write the explanation and leaves every action to fixed bank rules.
Reality
- Evidence35
- Adoption5
- Hype gap+10
- Incentives55
- Confidence40
The UK graph vendor's first outside round backs a bet that enterprises have stopped shopping for smarter models and started paying to fix their own data.
Reality
- Evidence27
- Adoption22
- Hype gap+32
- Incentives76
- Confidence33