Skip to content

Build1 publisher3 min readPublished

A docs bot that refuses to answer is working: the case for an evidence gate over a bigger window

An architecture decision record for a gaming moderation assistant puts a policy stage between retrieval and generation, and counts abstention with a reason code as a successful outcome.

The Engineer · Build desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

Illustration accompanying A docs bot that refuses to answer is working: the case for an evidence gate over a bigger window
Generated illustration

What happened

  • An architecture decision record published on dev.to concerns a docs chatbot that classifies gaming moderation reports before human review, and chooses evidence gating over a larger context window.
  • Retrieval may propose evidence, but a separate policy must decide whether the system may answer.
  • Embeddings answer a proximity question and do not prove that a retrieved passage governs this game mode, policy version, region, or report type.
  • Chunking can preserve more local meaning and a larger context window can carry more text, yet neither mechanism turns weak evidence into a warranted decision.
  • The chosen architecture is two-stage: retrieve candidate policy passages, then gate generation on evidence quality and scope.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

An architecture decision record published on dev.to for a gaming moderation assistant argues that the cure for a docs chatbot that invents answers is neither a larger context window nor better chunking, but a separate policy stage that decides whether the system is permitted to answer at all [1][2]. The operational consequence is the interesting part: a report that fails the gate goes to a human with a machine-readable reason, and that path is recorded as a success rather than a miss [6][7].

The reasoning against the usual knobs is mechanical. Embeddings answer a proximity question; they do not establish that a retrieved passage governs this game mode, policy version, region, or report type [3]. Chunking can preserve more local meaning and a bigger window can carry more text, but per the author neither mechanism converts weak evidence into a warranted decision [4]. A fluent category label can still be unsupported [11].

So the design is two-stage: retrieve candidate policy passages, then gate generation on evidence quality and scope [5]. Passing reports get a suggested moderation category plus citations [6]. Failing ones get routed out with reasons such as no_policy_match, scope_conflict, or ambiguous_evidence [7]. Three invariants define the boundary: the answer must cite text supporting the selected category, every cited passage must carry the policy version and scope used at retrieval time, and conflicting passages must not be silently averaged into a confident label [8]. The author's argument for stating them this way is testability - each can be checked before and after generation, which a prompt instruction to "use the context" cannot [9].

Generation stays outside the authority boundary. The model summarizes and suggests; the moderation service owns the final state transition, validates the response schema, and routes uncertain cases to reviewers [10]. The analogy offered is OTP delivery: a provider accepting a request does not establish that the user received the message, so each boundary needs its own observable result [12].

The taxonomy is where teams will get the most immediate value. The post splits "wrong" into four distinct failures - retrieval, scope, evidence, and generation - each with a different repair [13][14][15][16][17]. Calling all four hallucination is what leads a team to tune the retriever when the missing component is an authorization rule [18]. The concrete failure modes are unglamorous: a current policy retrieving an obsolete appendix with similar wording, a report that says harassment while describing impersonation, a transcript that drops the proper noun distinguishing a player from a game item, and a long context holding the correct paragraph and a contradictory one simultaneously [19].

Measurement follows the same split. Retrieval evaluation asks whether labeled support appears among candidates, gate evaluation asks whether supported cases pass and conflicting ones stop, and answer evaluation checks that each citation actually backs the claim [22]. End-to-end accuracy alone hides compensation, since a generator can guess correctly after retrieval has already failed [23].

What to watch: this is a design argument, not an evaluation report, and it carries no measured accuracy, abstention, or latency figures [24]. The author concedes the tradeoff explicitly, preferring quality over shaving latency from the happy path [25]. The numbers worth demanding from anyone shipping this pattern are the abstention rate, its distribution across the reason codes, and how often reviewers overturn a gate decision in each direction.

Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories