Skip to content

Build1 publisher3 min readPublished Updated

A NIST AI RMF-mapped RAG system for $25 a month plus a third of a cent per query

Lakshman Pandey's write-up maps a production retrieval system to all four NIST functions. The controls that exist are documents and evals; the ones that would gate a deploy are still marked future.

The Engineer · Build desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

Illustration accompanying A NIST AI RMF-mapped RAG system for $25 a month plus a third of a cent per query
Generated illustration

What happened

  • Lakshman Pandey published a dev.to article dated August 2026 documenting how a production retrieval-augmented generation system serving UK arts and culture clients implements NIST AI Risk Management Framework controls, with decisions, trade-offs and outcomes.
  • The system costs $0.003-0.005 per query and $25/month in infrastructure, with EU data residency and an eval framework to prevent quality degradation.
  • The NIST framework has four functions: GOVERN, MAP, MEASURE, MANAGE.
  • Stack: Streamlit Cloud frontend, Supabase pgvector vector DB in EU-West-2, Voyage AI embeddings at 1024 dimensions, Claude Haiku 4.5 via direct REST API, Langfuse observability, and an MCP server for Claude Desktop.
  • ADR-001 records the decision to use Supabase pgvector in EU-West-2 (London) because UK public-sector cultural clients require UK/EU data residency for GDPR compliance; the logged risk is vendor dependency on Supabase, and the mitigation is that the eval framework plus the ADR ensure reversibility.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

A developer has published a control-by-control mapping of a live retrieval-augmented generation system to the NIST AI Risk Management Framework, for a system serving UK arts and culture clients [1]. The line that matters to anyone being quoted an enterprise governance platform: the whole thing runs on $25 a month of infrastructure plus $0.003 to $0.005 per query [2].

The framework has four functions, GOVERN, MAP, MEASURE and MANAGE [3]. What fills them in this build is paperwork and instrumentation, not products. GOVERN is an architecture decision record, ADR-001, which records the choice of Supabase pgvector in EU-West-2 (London) because UK public-sector cultural clients require UK/EU data residency for GDPR, names the resulting vendor dependency as the risk, and cites the eval framework plus the ADR itself as the reversibility mitigation [5]. The stated policy is that user data stays in the EU, with calls to Claude and Voyage treated as transient, and Langfuse holding an audit trail of each query's origin and destination [6]. MAP is an inventory over a corpus of 108 documents, rated LOW-RISK on the grounds that the system is retrieval-grounded, the corpus is small and controlled, stakeholders are few, and no real-time safety-critical decisions are involved [7].

Note who signs that rating. Decision authority is described as a solo architect with client stakeholder approval loops, the client approves governance policies and validates output quality, and the operations role is listed as TBD [8]. This is self-assessment with a customer in the room, which is what most small vendors actually have.

MEASURE is the part worth copying. There is a Ragas-based suite of 18 golden questions scoring faithfulness, context precision, context recall and answer relevancy, chosen respectively to catch hallucination, irrelevant retrieved chunks, missed chunks, and off-question answers [9]. Langfuse traces every production query; the sample trace shows 450 input tokens, 85 output tokens, $0.0031 and 1250 ms [10]. Cost per query is broken out as $0.0005 for embedding and $0.003 for generation [11], so generation is roughly 97 percent of marginal cost [4]. The latency target is under two seconds against about 1.2 seconds observed, about 40 percent inside the target [12][5].

MANAGE is thinner, and the author does not pretend otherwise. Today it is a system prompt constraining answers to the reports and try-catch handling that logs errors to Langfuse [13]. Rate limiting and a monthly cost ceiling are marked future [14], as are a human-in-the-loop gate for sensitive queries, prompt caching claimed to cut costs 25 to 50 percent, Presidio PII redaction, and a CI/CD regression gate requiring the eval suite to pass before deploy [15]. Until that last item lands, the eval suite is a report, not a gate.

The economics explain the architecture. At $0.0031 a query, the $25 fixed monthly cost dominates until roughly 8,000 queries a month [2]; 10,000 queries costs about $56 all in [3], and the floor is $300 a year [1]. Managed Supabase was chosen for backups, EU residency and zero DevOps, at the price of lock-in and a moderate migration cost, on the reasoning that a team of one has no DevOps capacity and clients demand EU residency [16].

Watch whether the regression gate and the cost ceiling ship, and whether an operations owner is named. Those three are the difference between a documented posture and an enforced one.

Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories