Build1 publisher3 min readPublished
A citation graph lets a legal RAG agent reach the penalty for breaking Section 26
A developer's GraphRAG build turns Thailand's 96-section PDPA into a 190-node graph so an agent can walk from Section 26 to the fines that cite it. The published part argues the vector-search gap from vocabulary and reports no measured recall figures.
The Engineer · Build desk

What happened
- The test question asks whether a company may collect employees' health data and what the penalty is, and Section 26 answers the first half with the words health data.
- The stack is BGE-M3 dense and sparse embeddings in Qdrant, a NetworkX graph, a two-hop walk and a FastMCP server for a plain Python agent, all local in Docker.
- A second part of the series covers a bug that made the agent cite the wrong penalty section while every log looked fine.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- constraint Retrieval quality now depends on citation extraction: a reference the parser misses, including one written in Thai numerals, never becomes an edge, and no walk can reach the section behind it.
- decision Teams indexing GDPR face the same choice, since its higher fine tier is defined by a list of article numbers that a vector-only index is likely to keep apart from the duties it enforces.
- exposure Adding the graph adds a failure that logs do not show, so correctness checks have to compare the sections an agent cites against the statute text itself.
- cost Adoption needs no new service tier because the whole stack fits on one Docker host; the real cost is writing and testing the citation parser.
Embed the test question and the nearest chunk is Section 26, because it contains the words "health data" [4]. The fine sits in Sections 79 and 84. Their text reads "A data controller who violates Section 26 paragraph one or paragraph three ... shall be liable to an administrative fine of up to five million baht" [5]. That sentence mentions neither health nor employees [5]. The author says its vector lands far from the question's [6]. Section 84 is 58 sections after Section 26 [1]. "Statutes do not repeat themselves, they point," the author wrote [8].
The fix stores each citation as a directed edge: sec_84 -> REFERENCES_SECTION -> sec_26 [7]. Direction matters. The edge runs from the penalty to the prohibition, and vector search lands on the prohibition. A walk from Section 26 that followed only outgoing edges would never reach the fine. It has to follow incoming edges as well. The author's lawyer analogy describes exactly that: she "walks over to the sections they cite and the sections that cite them" [9]. The graph is sparse, at 380 edges over 190 nodes, or two per node on average [2]. A two-hop walk [3] over a graph that sparse should return a small neighbourhood, though a heavily cited section will have more neighbours than the average implies.
The rest of the stack is plain and well chosen. Chunks are cut at the section boundary, 96 chunks for 96 sections [3], so a citation and a chunk refer to the same unit. BGE-M3 writes dense and sparse vectors into Qdrant [3]. I'd expect the sparse side to help with literal tokens such as a section number. The graph lives in memory in NetworkX, and a FastMCP server exposes it to an agent loop written in plain Python, all on one Docker host [3]. For a graph of 190 nodes, an in-memory library is the correct size of database. The dual-level keyword scheme is borrowed from LightRAG [13].
On transfer: the author says every number comes from actually running the code [2]. They still describe one statute and one test question. The published part does not include a recall comparison between vector-only and graph retrieval. For the approach to carry to another law, the law has to cite by explicit number, a parser has to extract those numbers, and the answer has to sit within two hops. The PDPA data already uses Thai numerals [12], so the extractor handles a format a naive regex would miss. GDPR meets the first condition. According to the author, Article 83(5) defines the higher fine tier purely by listing the articles whose violation triggers it [10].
The graph brings its own failure mode. Part two of the series covers a bug that made the agent cite the wrong penalty section while every log looked fine [11]. Two penalty sections point back at one prohibition [5], so the walk hands the agent more than one candidate. The code is published as pdpa-graphrag-mcp [14].
What to watch
- Part two's explanation of the wrong-penalty-section bug, and whether the cause sits in the graph edges, the walk, or the agent's choice between Sections 79 and 84.
- A published recall comparison between vector-only and graph-assisted retrieval on the PDPA test questions.
- Whether the pdpa-graphrag-mcp citation extractor is reused on a statute such as GDPR, where fines are tiered by lists of article numbers.