Build1 distinct publisher3 min readPublished
A builder of voice intake systems argues the model is the least interesting part of a knowledge base, and that the first deliverable is a list saying which document governs each topic.
The Engineer · Build desk

Compiled by The EngineerSomething wrong?How this is made
Similarity is the only thing the pipeline measures. A question about leave entitlement matches the 2023 policy about as well as it matches the 2026 one, because both documents are about leave entitlement; nothing in either marks which is live [5], and a build that stuffs the top matches into a prompt has nowhere else to look [2]. The three years between the versions do not help [1]. The slide deck made for one client has the same problem in reverse: it reads like an official company position because that is how it was written [6]. Retrieval returns what it finds, faithfully [7].
The part worth arguing with is the claim that the model is the least interesting component [3]. The danger the post names is a model behaviour. A list of search hits lets a reader notice that the top result is a draft out of somebody's personal folder; a fluent summary of that same draft arrives in the identical register as a correct answer and hides the document underneath it [8]. The model did not create the contradiction. It removed the review step that used to catch it, which makes fluency an argument for curation rather than a replacement for it.
The two rejected access shortcuts fail at the same boundary. Filtering results after the fact and instructing the model to keep certain things secret both admit restricted passages into the context and then try to manage what happens next [17][18]. Pushing the check into the query instead means the index has to carry access rules from the source system, and group membership has to stay in sync as people change roles [19]. Curation gets described as the single highest-leverage day on the project [10], and the list it produces is three fields wide per topic [2], but nothing about that list stays true on its own. Document status changes and staff move teams without the index being told.
The clearest sign this is a governance job rather than an engineering one is the by-product. Teams doing the exercise routinely find that two departments have been operating on different rules for a year, which the author rates as worth more than the chatbot [11]. No retriever produces that finding. It comes from a person with authority reading the same procedure written three times by three departments that quietly disagree [6]. On the intake systems the author builds, including a voice and SMS product for Community Action Agencies answering in more than 100 languages, the work starts with eligibility rules and program policy rather than speech, and ambiguous policy is not rescued by a better model [15]. Engineers dislike this stage because there is no library for it [12], which is the honest reason it keeps getting deferred until it reappears as a retrieval bug.
Ranked by verification strength, evidence, and original report placement.
The dev.to post is written by a builder of AI voice agents and intake systems, who says that means building knowledge bases whether the client calls them that or not.
A RAG prototype takes an afternoon: chunk documents, embed them, stuff the top matches into a prompt, ship a chat box; pointed at the real drive, the same code starts lying with total confidence.
A prototype corpus is curated: good documents, current versions, one team, one format, and questions the builder could already verify. All of those conditions break at once when the pipeline is pointed at production content.
The real corpus has the 2023 policy and the 2026 policy side by side with nothing marking which one is live.
The real corpus also holds a slide deck made for one client that reads like an official company position, the same procedure written three times by three departments who quietly disagree, scanned PDFs, spreadsheets where the meaning lives in the layout, and a folder called Final Final.
Retrieval faithfully returns what it finds; if what it finds contradicts itself, the answer comes back confidently wrong in a voice indistinguishable from the answers that were right.
Follow any of these and your For You feed starts watching them — no settings page required.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
One practitioner essay, no data
The cluster is a single dev.to post by an interested practitioner. Its descriptive and architectural claims are internally coherent and mechanistically self-explaining, but nothing is quantified: no project count, no retrieval defect rates, no evaluation results, and no second source to corroborate or contradict. Several load-bearing generalisations are explicitly experience-based rather than measured.
One self-reported deployment
The only adoption signal is the author's own Fortell AI voice and SMS intake system for Community Action Agencies. It is a real shipped-system disclosure but self-reported, single-vendor, and unaccompanied by usage volumes; no other team, product or survey in the cluster shows the described governance-first practices being adopted.
Deflationary framing, thin proof
The piece argues against the prevailing emphasis on models, embeddings and chunking and makes no product, performance or capability promises, which pushes it slightly to the understated side. The offsetting factor is a set of confident superlatives - the model is the least interesting part, the highest-leverage day, the most common failure - presented without measurement, so the net gap is only mildly negative rather than strongly so.
Vendor author promoting build practice
The author sells the exact service the post describes - AI voice agents and intake systems - and names his team's product, Fortell AI, as supporting experience. The argument that success depends on business-side content governance and early requirements work aligns with billable discovery and consulting scope, and the post carries no disclosure beyond the self-description. This is ordinary practitioner marketing rather than concealed sponsorship, so the reading is elevated but not extreme.
Coherent but uncorroborated
Confidence is moderate-low: the cluster has one source, one author, and an evident commercial interest, so frequency and ranking claims cannot be checked. It is not lower because the technical prescriptions - query-time permission filtering, avoiding post-hoc result filtering, treating prompt instructions as non-enforcement - are mechanically self-consistent and would be recognisable to any reviewer, and because the source's own claims are clearly separated from its recommendations.
build
A docs bot that refuses to answer is working: the case for an evidence gate over a bigger window1 distinct publisher
build
Retrieval Is Not A Cheap Agent, And An Agent Is Not A Smart Retriever1 distinct publisher
build
The chunker sets the ceiling: how a character boundary turns a correct policy into a wrong answer1 distinct publisher
build
Don't start at the model layer: classify inputs by reliability, then let RAG wait1 distinct publisher
Distinct publishers with included, body-backed reporting in this cluster.
dev.to
1 article · August 25, 2026