Build1 publisher2 min readPublished
Kapa's Deep retriever edges a grep agent 0.65 to 0.61 on a benchmark it keeps private
Kapa says its Deep retriever scored 0.65 against 0.61 for a grep-running agent on 1,000 real company queries. Kapa built all seven systems and keeps the data private, so the result argues for its own design in a form buyers cannot reproduce.
The Engineer · Build desk

What happened
- Kapa says the grep agent took roughly five times as long as its Deep retriever to answer each query.
- Kapa's own Default mode tied the grep agent at 0.61 while answering in 3.3 seconds per query.
- Plain hybrid search scored 0.41, adding a reranker lifted it to 0.50, and query decomposition took it to 0.56.
- The rubric rewards enough context to answer, marks down unnecessary material and prefers better sources, such as a current Slack thread over an old internal document.
- Agents labelled the 1,000 cases after tuning against 170 human-marked cases, and Kapa did not report a numerical agreement rate.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- constraint The grep agent's 0.61 describes grep over documents as Kapa's pipeline ingested them, so a team grepping its own unprocessed files would need its own measurement.
- decision A 0.04-point lead can only settle a purchase once the labelling agents' error is known to be smaller than that, so buyers have to test on their own queries first.
- cost The answering model pays in time and context for every unnecessary chunk retrieval hands it, so a retriever's token volume is a running cost separate from its score.
Read as increments, the stage scores show where Kapa's gains come from. The reranker added 0.09, query decomposition 0.06, the move to Default 0.05, and Deep a further 0.04 [4]. Each stage adds less than the one before it. Publishing the score at every stage is good engineering practice. It shows that the reranker is the largest single step [4].
Kapa's score also measures more than whether the right document was found. The rubric marks down unnecessary material [6], so the grep agent's 40,000 tokens [10] count against it on the same scale that ranks it. Deep returned roughly 5,000 tokens [11], an eighth as much [1]. Some of the grep agent's deficit is probably that volume, measured a second way. In my view that is the right rubric for agents that hand their context to another model. It is also the rubric Kapa's design is built to win. Kapa argues that an agent which searches iteratively and prunes irrelevant documents returns a more useful result without making each query slower or more expensive to consume [14].
The ambiguity rule pulls the other way. When a query has several reasonable readings, retrieval must return material for each one and leave the agent to choose [7]. A system that prunes too hard loses points on those cases, so the rubric tests coverage as well as precision.
Speed is the weaker claim. Kapa reports 13 to 17 seconds per query for the grep agent [10] against about five for Deep [2]. That is a ratio of 2.6 to 3.4 [2], short of the five-times figure Kapa gives [4]. A ratio near five appears only when the slowest grep runs are set against Default's 3.3 seconds [5]. RuntimeWire's account of the post does not reconcile the two figures.
Whether the result transfers depends on the questions. Kapa's set mixes developer questions about documentation, APIs, code and GitHub issues with employee questions drawing on Slack, Confluence, Notion and Google Drive. It also covers support work on tickets, help-center articles and internal handbooks [8]. The cases come from Kapa's production traffic [1]. RuntimeWire notes that this may reflect Kapa's customers better than a synthetic set would, while the private method makes it harder for outsiders to judge how representative the cases are [18].
Kapa says production data is the reason the benchmark cannot be public [13]. That is a plausible reason. It also means the only runnable copy of the comparison [5] sits with a company that sells a shared retrieval layer for agents searching a company's scattered knowledge [16].
What to watch
- Whether Kapa publishes the agreement rate between its labelling agents and the 170 human-marked cases.
- An independent rerun of a grep agent against a tuned retriever on a public corpus, with token counts and latency reported.
- A clarification from Kapa reconciling the five-times latency claim with its reported 13-to-17-second grep times.