Skip to content

Build1 publisher3 min readPublished Updated

Semantic code search over a monorepo is now a plumbing job, and the plumbing is the hard part

A dev.to walkthrough wires chokidar, tree-sitter, Qdrant and an MCP server into a code search stack. Every part is off the shelf; the unresolved work is reindexing and retrieval quality.

The Engineer · Build desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

Illustration accompanying Semantic code search over a monorepo is now a plumbing job, and the plumbing is the hard part
Generated illustration

What happened

  • A tutorial titled "Build a Codebase Intelligence Tool Like repowise With a RAG-Assisted MCP for Your Monorepo" was published on dev.to, noting it was originally published on tamiz.pro, and describes building a production-grade RAG-assisted Model Context Protocol server that turns a local codebase into a queryable knowledge base for LLMs and CLI tools, framed as building your own version of tools like repowise.
  • The tutorial's architecture diagram shows a File Watcher (chokidar) feeding an Indexer (LangChain) feeding a Vector Store (Qdrant), with a Retriever (Hybrid) feeding an MCP Server (FastMCP) that serves IDE/CLI clients including Claude and Cursor.
  • The tutorial uses tree-sitter for language-agnostic parsing, stating this preserves semantic boundaries (functions, classes) instead of arbitrary character splits, and that the indexer splits code into AST-aware chunks, embeds them and stores them in a local vector database.
  • The sample chunker emits a chunk when a tree-sitter node type is one of function_declaration, class_declaration or method_definition, and sets the chunk id to `${filePath}:${node.startPosition.row}-${node.endPosition.row}`.
  • The embedding server is a local FastAPI app using sentence-transformers with the model BAAI/bge-small-en-v1.5 and normalize_embeddings=True, exposing a POST /embed endpoint on port 8000, recommended for privacy and zero API cost.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

A tutorial published on dev.to, originally on the author's own site, lays out a full reference build for a "RAG-assisted" Model Context Protocol server that turns a local monorepo into a queryable index, presented as a do-it-yourself version of tools like repowise [1]. What makes it worth reading is not any modelling idea but the parts list: every box in the diagram is a named, installable dependency, which moves the risk out of research and into integration.

The shape is three stages plus a serving path. A chokidar file watcher feeds a LangChain-based indexer, which writes into a Qdrant vector store; a hybrid retriever sits behind a FastMCP server that IDE and CLI clients such as Claude and Cursor consume [2]. Chunking is done with tree-sitter so splits land on functions and classes rather than arbitrary character offsets [3], keyed on `function_declaration`, `class_declaration` and `method_definition` nodes, with each chunk given an id of the form `path:startRow-endRow` [4]. Embeddings come from a local FastAPI service running sentence-transformers with `BAAI/bge-small-en-v1.5`, normalised, on port 8000, justified on privacy and zero API cost [5]. Prerequisites are Node 20+, Python 3.11+, a running Qdrant, and a repo under a million files [6].

Note where the model sits: behind an HTTP POST to `/embed` [5]. That boundary is the whole argument. Swapping the embedding model is a config change; the things you cannot swap are the chunker, the watcher's update semantics, and the ranking.

The watcher is described as updating the store incrementally [7], and that is where the id scheme starts to matter. Because ids encode start and end lines [4], inserting a line near the top of a file changes the id of every chunk below it, so an upsert-only path leaves the previous generation of chunks resident and still eligible for similarity search unless the indexer deletes by file path first [8]. The excerpt supplied stops mid-function inside `indexFile`, before any of that logic appears [9], so on the evidence available the hardest part of the build is the part not shown.

Retrieval quality has the same gap. Combining vector similarity with BM25 lexical search and AST context [10] is the right shape for code, where identifiers want exact matching and intent wants embeddings. But the material specifies no fusion weighting, no evaluation set, and no measured recall, latency or index size [11]. At the tutorial's stated ceiling of a million files [6], the distinction between a correct answer to "where do we check user permissions" and a plausible one is entirely ranking [12], and ranking regressions are silent unless you have fixed queries with known answers.

Watch three things if you build this. Whether your incremental path handles deletes and renames, not just edits [8]. Whether the exclude-pattern and language config, which the shared types expose as `excludePatterns` and `languages` [13], is doing enough to keep generated code and vendored trees out of the index. And whether anyone on the team owns a retrieval eval; the tutorial reserves a section for production hardening and an FAQ [14], but hardening a pipeline is not the same as knowing it answers correctly.

Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories