Skip to content

Build1 publisher3 min readPublished

The model was never the bottleneck: inside a Kannada RAG rebuild that blames retrieval

A scanned 346-page Kannada novel broke pure vector search. Hybrid BM25 plus dense retrieval with RRF, and a regex page router, took reported faithfulness to 0.92.

The Engineer · Build desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened

  • The author built a RAG agent for Heli Hogu Kaarana, a Kannada novel by Ravi Belagere, digitized from a scanned 346-page PDF, and published the ingestion pipeline, retrieval architecture and evaluation numbers.
  • The same LLM sat behind both v1 and v2; the author states the difference was entirely in retrieval architecture, and that the LLM was never the bottleneck.
  • V1 architecture was chunk, embed with multilingual MiniLM, store in ChromaDB, retrieve top-5, prompt Gemini.
  • The author writes that v1 demoed well but evaluated terribly.
  • The author stopped trusting the system when it confidently answered a question about page 120 with a passage from an entirely different chapter, with no citation and full confidence.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

A developer has published a rebuild of a retrieval-augmented generation app built over a scanned 346-page Kannada novel, Heli Hogu Kaarana by Ravi Belagere, along with the architecture and the evaluation numbers [1]. The argument worth taking seriously is the boring one: the same LLM sat behind both the failing and the working version, and the difference was entirely in retrieval architecture [2].

Version 1 was the default recipe: chunk, embed with multilingual MiniLM, store in ChromaDB, retrieve top-5, prompt Gemini [3]. It demoed well and evaluated terribly [4]. The tell was a question about page 120 answered confidently with a passage from a different chapter [5].

Three properties of the corpus explain the failure. The book existed only as a scanned physical object, and Kannada OCR is hard because of ligatures, conjunct characters and noisy print [6]. Kannada is agglutinative, so a name like Himavant appears in several inflected forms depending on grammatical case, and the author's contention is that multilingual embeddings trained mostly on high-resource data compress those forms into a weak, inconsistent semantic space [7]. Literary text is sparse and specific, and rare colloquialisms, proper nouns and page references are precisely where cosine similarity fails [8]. Under dense-only retrieval, proper nouns vanished and page queries returned whatever felt close [9]. Generic models have essentially never seen this book, so with bad retrieval they invent rather than retrieve [10].

The fix has two load-bearing parts. Sparse BM25 runs alongside dense retrieval and the two ranked lists are fused with Reciprocal Rank Fusion at k = 60, followed by a cross-encoder rerank with a confidence guardrail, then Gemini with a Groq fallback [11][12]. Separately, a deterministic regex router classifies exact-page queries and bypasses semantic search entirely, resolving them through a metadata lookup in about milliseconds, which the author reports as 100% precision with zero hallucination surface [13]. That precision is a property of the router's classifier, not of the index: a page query the regex fails to recognise falls back to the semantic path, where the old failure mode still lives [14]. The stated design principle is that no single retrieval signal makes the final decision [15].

Ingestion carries the rest of the weight: pdf2image at 300 DPI, OpenCV denoising and thresholding, Surya OCR, indic-nlp normalization, and semantic chunking with page metadata attached to every chunk [16]. The author calls the metadata the quiet hero, because it is what makes citation possible and what the deterministic router runs on [17].

Reported results are 0.92 RAGAS faithfulness and 0.89 context recall on a 50-query golden set, at 2.8s P50 end to end on serverless [18]. Read those with the denominator in view: on 50 queries, a single query is worth 0.02, so the third digit is noise [19]. This is also one self-reported build, and the supplied write-up gives no v1 RAGAS scores to compare against, only the qualitative claim that v1 evaluated badly [20].

What to watch: whether OCR quality caps the whole system, since the author's own rule is that retrieval is only as good as ingestion [21]; whether the regex router's coverage holds as query phrasing drifts away from the patterns it was written for [14]; and whether the confidence guardrail actually refuses, or just passes low-confidence context through to a fluent answer [12].

Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories