Leadership1 distinct publisher3 min readPublished
Genie answers engineers in Slack rather than customers in an app. The 45,000-questions-a-month figure Uber published gives other planners a baseline they rarely get, even without a disclosed deflection rate.
The Board Room · Leadership desk

leadership
Uber's security experts withheld the Slack channels until Genie's retrieval improved1 distinct publisher
leadership
Uber is retrofitting attribution onto the agent platform it already shipped1 distinct publisher
product
Canva's forecast cut turns model routing into a product line item1 distinct publisher
product
ChatGPT Work's real ask is your Slack, and somebody has to say yes on everyone's behalf1 distinct publisher
Compiled by The Board RoomSomething wrong?How this is made
Volume is the part of this disclosure that transfers to other companies. About 45,000 questions a month divides to roughly 1,500 a calendar day, or about 2,100 per working day on a 21-day month [12]. If that traffic is spread across the hundreds of channels Uber says carry on-call demand [3], and "hundreds" means somewhere between 200 and 500, each channel is seeing perhaps 90 to 225 questions a month, four to eleven per working day [14]. That is the shape of a problem that rarely gets funded on its own: no single team's queue looks bad enough to justify a platform, while the aggregate is a standing charge against engineering time [2].
Internal support pays back for reasons that have little to do with model quality. The corpus belongs to the employer, and the audience can usually tell a wrong answer from a right one. The cost of a miss is a wasted minute rather than a commitment made to a customer. Uber describes what it faced as a findability problem rather than a generation problem [3], and its preference for retrieval-augmented generation on time-to-market grounds [4] fits a body of documents that changes faster than any training cycle would.
The board-deck version of this reads "AI now handles 45,000 questions a month." The post does not claim that. It gives the channel volume, the architecture of scraped wiki and internal Stack Overflow content embedded with an OpenAI model into a vector database and retrieved into an LLM prompt [5], and three hallucination controls: relevant retrieval, verification against authoritative sources, and keeping the index current [7]. There is no deflection rate, no accuracy measure and no figure for engineer hours returned [13]. The disclosure is a denominator, and anyone building a business case from it supplies the numerator themselves.
The obvious objection is that a copilot sitting over fragmented documentation just launders weak docs into fluent answers. Uber's own controls concede the mechanism, since two of the three depend entirely on the authority and freshness of the sources behind them [7], and the security screen narrows the corpus further, because Uber says many internal data sources cannot be exposed in Slack channels at all [6]. Even so, that caps the ceiling rather than voiding the case. Retrieval quality is bounded by document quality, which turns a documentation problem that was previously ignored into one with a visible score.
The sequencing worth noting is what the retrieval choice moves onto next quarter's budget. Fine-tuning would have bought a model refresh cadence; retrieval buys an ingestion pipeline, and Uber says it runs a custom Apache Spark application to load the vector store [8]. That is a data engineering commitment with an owner and a maintenance schedule, alongside the feedback loop Uber built for users to rate responses [11]. A team approving a copilot like this in one quarter is approving the person who maintains its index in the next, and the question volume is the only part of the case that arrives with a number already attached.
Ranked by verification strength, evidence, and original report placement.
Genie scrapes internal data sources including Uber's internal wiki, internal Stack Overflow and engineering requirement documents, creates vectors using an OpenAI embedding model, stores them in a vector database, converts each Slack question into embeddings, searches for relevant embeddings, and uses the results as prompts to the LLM.
Uber says a custom Apache Spark application performs the ETL steps for ingesting data into the vector database.
Uber says teams such as the Michelangelo team run Slack support channels where internal users ask for help, and that people ask around 45,000 questions on these channels each month.
Uber says high question volumes and long response wait times reduce productivity for both users and on-call engineers.
Uber says many questions could be answered from existing documentation, but the information is fragmented across its internal wiki (Engwiki), internal Stack Overflow and other locations, so users often ask the same questions repeatedly, creating high demand for on-call support across hundreds of Slack channels.
Uber chose retrieval-augmented generation rather than fine-tuning an LLM for Genie, saying fine-tuning requires curated high-quality diverse examples plus compute to keep the model updated, while RAG needed no such examples and reduced time to market.
Distinct publishers with included, body-backed reporting in this cluster.
1 article · August 27, 2026
Follow any of these and your For You feed starts watching them — no settings page required.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Detailed first-party architecture account, no outcome data
The single source is Uber's own engineering blog and is unusually specific about mechanism - RAG choice rationale, Spark ETL stages, langchain chunking, OpenAI embeddings, Terrablob and Sia - which supports the descriptive claims well. But there is no independent corroboration, no evaluation of answer quality, and no reported deflection, resolution or time-saved metric, so any claim about Genie working is asserted rather than evidenced.
One internal deployment, scope undisclosed
Adoption evidence is limited to Uber's own statement that it built and runs Genie against Slack support channels, plus the 45,000-questions-a-month load it targets. The disclosed volume measures the problem, not Genie's traffic; the post gives no count of channels or teams served, no user numbers and no feedback-loop results, and there is no third-party or external adoption of the approach in the supplied material.
Mildly overstated: capability framing without efficacy data
The post's language is comparatively restrained for a vendor-adjacent blog and it does not claim a productivity win in numbers. Still, it presents hallucination as 'solved' through three general measures and frames Genie as a copilot that optimizes on-call question-answering while publishing zero accuracy, deflection or time-saved evidence, so the capability framing runs modestly ahead of what is shown.
First-party engineering-brand publication
The sole publisher is the company that built the system, publishing on its own engineering blog, where the incentive is to showcase platform sophistication and recruiting appeal rather than to report shortcomings. No product is being sold and no pricing or customer pitch appears, which limits commercial incentive, but selective disclosure - detailed architecture, omitted outcomes - is consistent with that publishing motive.
Descriptive facts solid, outcomes unknown
Confidence is high that Uber said what the claims report and that the described architecture and 45,000-questions figure are as published, since the source is direct and specific. Confidence is low on anything about effectiveness, scale of rollout or durability of the hallucination controls, and the single-publisher cluster leaves no cross-check.