Skip to content

Build1 publisher3 min readPublished

A RAG Pipeline in 200 Lines of TypeScript, and the Parts the Frameworks Hide

A dev.to writeup drops LangChain, LlamaIndex, hosted vector stores and the cloud model, leaving six files, seven npm packages, and one cloud dependency still in place.

The Engineer · Build desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

Illustration accompanying A RAG Pipeline in 200 Lines of TypeScript, and the Parts the Frameworks Hide
Generated illustration

What happened

  • A dev.to post by an author publishing as Apurwa describes building a RAG pipeline in six files and a bit over 200 lines of TypeScript, with nothing imported the author cannot explain.
  • The build uses no LangChain, no LlamaIndex, no hosted vector database and no cloud LLM; the model runs on the author's laptop.
  • The author read four RAG tutorials and still could not have said what an embedding actually was, why cosine similarity was the metric everyone used, or what would happen if the documents were 800 pages instead of 8; the author could copy the code but not debug it.
  • The tutorials the author found consisted of about twelve lines using RetrievalQAChain.fromLLM with a vector store retriever, a Pinecone key, a screenshot of it answering one question about one PDF, and a confident closing paragraph about production readiness.
  • The pipeline runs in five steps: extract raw text from the PDF, chunk it into overlapping windows, embed each chunk as a list of numbers, embed the question and retrieve the closest chunks, then paste those chunks into a prompt and send it to a model.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

A developer writing as Apurwa on dev.to published a retrieval-augmented generation pipeline built by hand: six files, a bit over 200 lines of TypeScript, no LangChain, no LlamaIndex, no hosted vector database, and the language model running on the author's own laptop [1][2]. The interesting part is not the line count but the stated reason for it: after reading four framework tutorials, the author could copy the code and could not debug it [3].

The tutorial pattern being rejected is specific and recognisable. Twelve lines built on `RetrievalQAChain.fromLLM`, a Pinecone key, a screenshot of the thing answering one question about one PDF, and a closing paragraph about production readiness [4]. What those twelve lines did not explain, according to the author, was what an embedding actually is, why cosine similarity is the metric everyone reaches for, and what happens when the documents are 800 pages instead of 8 [3]. The third question is the operational one, and it is the one a single-PDF screenshot cannot answer.

Stripped down, the pipeline is five stages: extract text from the PDF, chunk it into overlapping windows, embed each chunk as numbers, embed the question and find the nearest chunks, then paste those chunks into a prompt [5]. The framing is blunt: models cannot read your files, so find the relevant paragraphs yourself [6]. Keyword search will not do it, because someone asking how to stop duplicate rows never types DISTINCT [7].

The local-model choice does real work. LM Studio serves any instruct-tuned model that fits your RAM on `localhost:1234` in the OpenAI API format, so the official `openai` SDK talks to it unmodified, with no API key, no per-token cost, and no internet [8]. The author's argument is that this removes the billing dashboard from the loop and forces you to hit a small context window early, which is where one of the more instructive bugs came from [9]. That is the correct trade to make while learning: fail against the constraint you will ship into.

The "no cloud" framing has one seam. Embeddings still go out to Voyage AI, on a free tier of 200M tokens a month, with the key read implicitly from `VOYAGE_API_KEY` so you never pass it explicitly [10]. The model is local; the pipeline is not offline [3]. And a client that picks up credentials from the environment on its own is the same species of convenience the post objects to elsewhere, just cheaper to unwind.

The dependency list is four runtime packages and three dev packages, seven in total [11][1], which is the honest cost of "no framework." The code that ships is roughly 33 lines per file on average [2], and the extraction file is where the reasoning is densest: `fs.readFileSync` with no encoding argument returns a Buffer because a PDF is binary and `"utf-8"` would hand you mangled text [12]. The synchronous read inside an `async` function blocks the event loop, which is invisible for a CLI that reads one file at startup and would not be for a server handling concurrent uploads, where `fs/promises` is the swap [13]. `pdf-parse` is pinned at 1.1.1 on purpose [14].

Watch two things. The excerpt available here stops at the extraction stage, so the chunking, embedding, retrieval and generation code, plus the four bugs and the debugging method the post promises, are not in it [15][16]. And watch the pinned parser version: a deliberate pin is a bug someone already paid for.

Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories