Build1 publisher3 min readPublished
Retrieval Is Not A Cheap Agent, And An Agent Is Not A Smart Retriever
A dev.to post argues the "RAG or agents" question is malformed. The interesting part is the invoice: one axis out of five justifies the premium, and most buyers never use it.
The Engineer · Build desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction
What happened
- A developer writing on dev.to says a client emails the same question at least once a month: "Should we use RAG or agents?"
- The post argues the phrasing is wrong because RAG and agents are not two versions of the same thing: RAG answers questions from your knowledge base, an agent takes actions toward a goal.
- The post states that conflating the two is how you end up with a $3,000/month "agent" that should have been a $0.20 RAG pipeline, or a RAG pipeline that hallucinated its way into taking an action nobody authorized.
- The post's one-sentence summary: RAG controls what the model knows; agents control what the model does, and everything else including cost, latency and failure modes follows from that single difference.
- The RAG flow described is short and deterministic: user question, embed, vector search, top-k chunks, prompt, answer, with documents stored as embeddings in a vector database.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
A developer writing on dev.to reports getting the same client email at least once a month: should we use RAG or agents? The post's answer is that the question is malformed, because the two are not competing implementations of one capability [s1c1][s1c2]. The reason an operator should care is the invoice attached to the confusion: by that account, conflating them is how you end up with a $3,000-a-month "agent" doing work a $0.20 retrieval pipeline would have done, or a retrieval pipeline that hallucinated its way into taking an action nobody authorised [s1c3].
The distinction is short enough to keep. Retrieval governs what the model knows; agency governs what the model does, and cost, latency and failure modes all follow from that single difference [s1c4].
Retrieval is a straight line: question, embed, vector search, top-k chunks, prompt, answer [s1c5]. The post prices it at $0.01 to $0.05 per query, one round trip, one prompt to reason about, with every answer traceable to a chunk that actually exists [s1c6][s1c7]. Its weakness is stated without hedging: it cannot act. It can tell you the refund policy and cannot issue the refund; it can summarise a log and cannot rotate the credential it just found [s1c8].
An agent is a loop rather than a call: observe, reason, call a tool, use the result, repeat until the goal is met, a budget is hit, or it escalates [s1c9]. The defining property is multi-step action without a human authoring each step in advance [s1c10]. Its default knowledge is training data plus whatever a tool hands back, so an agent with no retriever attached will answer questions about your private policy from vibes, and then act on them [s1c11]. Which is why, the post says, most production agents quietly contain a retrieval pipeline inside [s1c12].
The scoring is where the budget argument lands. On answer quality over your own content both score 4, because the agent's grounding comes from the retriever it contains; without one it is a 1 and confidently wrong, making this a dependency rather than a contest [s1c13]. On action capability retrieval scores 1 and the agent 5 [s1c14]. On cost per task retrieval scores 5 and the agent 2, with a five-step task and two tool calls running roughly 5 to 10 times a single generation [s1c15]. On latency, 5 against 3: a second or two versus seconds to minutes, which a user watching a chat window feels immediately [s1c16]. On predictability and governance, 4 against 3 [s1c17]. Added up, that is 19 out of 25 for retrieval against 17 for the agent, and the entire agent advantage sits in one column [1].
Put the money next to it. At the post's own $0.01 to $0.05 per query, a $3,000 monthly agent budget buys between 60,000 and 300,000 retrieval queries [2]. The same multiplier implies something like $0.05 to $0.50 per agent task [3]. Neither number is an argument against agents; they are an argument against buying action capability for a system whose entire job is to answer questions, which the post frames as paying a heavy premium for a capability you never use [s1c15].
The governance asymmetry is the part that does not show up in a pilot. Retrieval's failure is a wrong sentence; an agent's failure is a wrong action, which is why actions need budgets, permission layers and human approval [s1c18].
Two things worth checking this week. First, whether anything you call an agent has a retriever attached, because if it does not, its confident answers about your own documents are unsourced [s1c11]. Second, whether the permission scope of any deployed pipeline was reviewed by someone who understood it could mutate state rather than merely describe it [s1c18].