Build1 publisher3 min readPublished
Inference outside the Django app is what makes a multi-tenant RAG support product operable
A solo builder's writeup of an AI support SaaS puts the model behind a service boundary. Isolation, per-answer cost and freshness turn out to be architecture problems, not prompt problems.
The Engineer · Build desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction
What happened
- The author describes building AI-Autofy over several months, an AI customer-support SaaS designed to let businesses train an assistant on their own website, documents, FAQs and business data.
- The web application is built with Python and Django.
- Django handles customer accounts, subscriptions, AI configuration, knowledge-base management, chat history, analytics, widget configuration, tenant separation and integrations.
- The AI inference layer is separated from the main Django application; the web application does not run the language model itself, and requests are sent to an AI service responsible for generating responses.
- A typical Django request might take milliseconds, while an AI response may involve retrieval, prompt construction, GPU inference, streaming tokens, tool calls and live-data lookups.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
A developer writing on dev.to described the architecture of AI-Autofy, an AI customer-support SaaS built over several months to let businesses train an assistant on their own website, documents, FAQs and business data [1]. The load-bearing decision in the writeup is not the model choice: the inference layer sits outside the Django application, which never runs the language model itself and instead calls a separate AI service to generate responses [2][4].
Django keeps the stateful, boring surface: accounts, subscriptions, AI configuration, knowledge-base management, chat history, analytics, widget configuration, tenant separation and integrations [3]. The author's stated reason for the split is that the two workloads have nothing in common operationally. A typical Django request may take milliseconds, while an answer can involve retrieval, prompt construction, GPU inference, token streaming, tool calls and live-data lookups [5]. Separating them lets the web application and the AI infrastructure scale independently [6], and lets the model change without redesigning the SaaS application [7]. That second consequence is the one that matters commercially: the model becomes a replaceable component rather than a dependency baked into the product.
Isolation is the requirement that makes the boundary more than tidiness. The retrieval path is question, embedding, search of the business knowledge, retrieved documents, prompt construction, grounded answer [8]. The author's rule is blunt: a document belonging to Company A must never appear in a response generated for Company B, so every retrieval request has to remain scoped to the current tenant [9]. Tenant filtering is listed alongside chunk size, metadata, similarity thresholds, query rewriting, the number of retrieved chunks and prompt construction as the variables that determine retrieval quality [10]. In practice that makes the tenant filter a retrieval parameter passed on every call, not an assumption inherited from web request state.
Predictable inference cost appears as one of the things that turn a chat widget into a system, along with reliable business-specific answers, tenant isolation, live data, product information, images and analytics [19]. The only cost mechanism actually described is relevance gating. If product data is injected into every request, the model may start mentioning products when the customer asked about opening hours; the fix is a relevance check before enrichment, sketched as `if is_product_question(message): context += get_product_data()` [16]. The author reports this improves both response quality and token efficiency [17], and separately that relevance filtering was one of the most important parts of the system, since too little context produces incomplete answers and too much fills the context window with irrelevant material [11][12].
Freshness gets the same structural treatment. A vector database suits relatively static content, so website text, FAQs, documentation and policies are handled as knowledge data, while products, prices, availability and business-system information are treated as live data, because embedding yesterday's catalogue is not necessarily the right answer to a question about current availability [13][14]. The hard part, per the post, is deciding when live data should be queried at all, since a catalogue should not enter a prompt merely because it exists [15].
Two things would move this from design argument to result. First, how the gate is implemented: if `is_product_question` is itself a model call, part of the token saving goes back out the door, and the post does not say [16]. Second, where tenant scope is enforced. The post states the requirement [9] but not the enforcement point, and a tenant ID trusted from the caller is a contract rather than a boundary. Note also that the writeup carries no latency, cost or accuracy figures [20], and the discussion of dynamic images, which the author says has the same relevance problem, is not developed in the material available [18][21].