Build1 publisher3 min readPublished
98ms repo maps: what moving symbol indexing out of Node actually buys an agent
LiuHe's own numbers put a full index of 1,482 files at 9.7s and every repo map after that at 98ms. The more interesting figure is the crash budget it replaced.
The Engineer · Build desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened
- LiuHe is a code-operation toolchain designed for LLMs rather than humans, combining Node orchestration, a Rust tree-sitter parse daemon, a SQLite symbol index and transactional journal-backed writes; it is MIT-licensed, at v0.4.6, with zero-build deploy (no cargo, no npm install).
- The writeup is titled and framed as cutting repo-wide symbol indexing for LLM agents from 30s to 98ms; repo_map returns in 98ms, where it previously took tens of seconds, via a Rust tree-sitter parse daemon plus SQLite index plus incremental self-heal.
- The first version answered "what symbols exist in this repo?" by walking the tree and parsing every file on every request; on a 347-file project that took tens of seconds, and on a real Ansible repo of 1,482 files it was worse.
- Agents ask for repo maps constantly, because every tool call needs file to symbol to reference context.
- All CPU-bound AST work lives in a Rust parse daemon (tree-sitter plus tokio plus rayon) that talks to Node over a Unix socket, with zero-copy source slicing, no per-node N-API boundary crossings and true parallelism.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
The maintainers of LiuHe, an MIT-licensed code toolchain built for LLM agents rather than people, have published an architecture writeup saying repo-wide symbol lookup now returns in 98ms, down from tens of seconds, after CPU-bound parsing moved out of Node into a Rust tree-sitter daemon fronted by a SQLite index [1][2]. The number matters less than where the cost was sitting: according to the writeup, every tool call needs file to symbol to reference context, so index latency is charged against each step of an agent loop, not once per session [4].
The first version answered "what symbols exist in this repo" by walking the tree and parsing every file on every request [3]. On a 347-file project that was tens of seconds; on a 1,482-file Ansible repo it was worse [3]. The headline framing of 30s to 98ms works out to roughly a 300x cut [18], but the honest way to read it is that the old design was doing full work for a cached question.
Three changes, in order of consequence. All AST work moved into Rust (tree-sitter with tokio and rayon), speaking to Node over a Unix socket, with zero-copy source slicing and no per-node N-API crossings [5]. Parse results land in SQLite in WAL mode, per workspace, so later queries are point lookups instead of re-parses [6]. Freshness is handled by mtime plus dirty flags, and if the sha256 of the Rust binary changes the whole database is marked dirty and rebuilt [7]. The reported result: a full index of 1,482 files in 9.7s, or 153 files per second, and 98ms repo maps thereafter [8]. A cold build therefore costs about the same as 99 warm map calls [19], which is the part to check against your own upgrade cadence, since a binary bump triggers that rebuild by design [7].
The stability story is arguably the bigger operational win. The team reports the in-process tree-sitter binding died on parse exceptions, taking the whole MCP server with it, on average every two to four hours of use, with GC pauses and JS-to-C crossings stalling batch indexing [9]. In the daemon, catch_unwind converts a parser panic into a PARSE_PANIC error code and the server survives [10], and the parse path has no GC pauses [11].
The write side follows the same logic. edit_transaction is all-or-nothing with an undo journal, and the team says it tested kill -9 mid-write: the half-written transaction rolled back and source files were untouched [12]. Writes are version-anchored with optimistic concurrency, so a stale write fails loudly rather than silently overwriting [14], and errors return a suggestion and a next_action the model can reissue verbatim [13]. The wider surface is 44 tools across read, analyze, edit, gate, verify and system [15], with the gate family doing zero LLM calls [16].
Every figure here is self-reported by the project, which says the benchmarks are reproducible from its benchmarks/ directory [17]. Worth watching: whether 153 files per second holds on larger, mixed-language trees; whether sha256-triggered full rebuilds become the new tax for teams that update often; and whether the undo journal survives failures less clean than kill -9.