Build1 distinct publisher3 min readUpdated
Bruna Pereira's SafeChat talk puts a number on unsafe chat traffic before it puts a model in the message path. That ordering, not the model, is what survives a rewrite.
The Engineer · Build desk

Compiled by The EngineerSomething wrong?How this is made
Nine percent is a generous ceiling on "small single digit", and it is enough to do the architecture arithmetic. Of 4 million daily messages, at most 360,000 would ever need an expensive opinion, and at least 3.64 million would not [15]. At the 2 to 10 second latencies Pereira quotes for her team's LLM calls [7], that escalated slice works out to roughly 8 to 42 model calls in flight at any moment, averaged across the day [16]. Peaks will be worse than the average, but that is a fleet you can reason about, and it is why the cheap layer is not a compromise: it turns an unaffordable per-message problem into a small one you can afford to do well.
The latency arithmetic is less forgiving. A sub-second decision budget against 2 to 10 seconds of measured model latency is over budget by 2x at best and 10x at worst [17], and no gain in classification quality returns any of that time. Model selection is not the lever here. Routing is.
Which is what makes the rewrite the more useful half of the talk. The published transcript stops as the small classifier's three jobs are introduced and never lists them [19], so the replacement's internals are not in this record. What Pereira says she wants the audience to keep is a pattern rather than an implementation, one she claims applies to almost any AI use case [3]. On the evidence she has already laid out, that pattern is an ordering: get the base rate off your own traffic first, then size a cheap filter to it and let the expensive model see only what the filter cannot settle. None of that is bound to a particular model, which is exactly why a system that worked could be thrown out [2] without taking the design with it.
The bill for the pattern lands in one place. "Right almost all of the time" [11] on a safety product means the remainder is an abusive message delivered, to two parties whose relationship runs somewhere between six and forty minutes and who have no history to fall back on [13]. The talk gives no false-negative rate for the cheap layer. Against a stated goal that every message reaching its destination has been classified as safe [5], that is the figure worth asking for, and it is a decision about tolerable residual risk rather than a modelling detail. Verbal abuse is described as a meaningful share of incidents on the platform, and feeling safe is treated as a product metric on equal footing with being safe [14], which means the misses are measured in both.
Follow any of these and your For You feed starts watching them — no settings page required.
Ranked by verification strength, evidence, and original report placement.
Pereira said the team built a system to solve a real safety problem in production, and when it worked, they threw it all away to build something even more powerful.
The analysis confirmed most messages are safe and put a number on it: only a small single-digit percent of messages were unsafe.
Pereira said that number shaped the entire architecture, because it told the team it could build an aggressive cheap layer that was right almost all of the time.
Bruna Pereira is a software engineer at DoorDash who leads the trust and safety engineering team from the company's hub in Sao Paulo, Brazil, and presented SafeChat: Building AI-Powered Safety Systems at Scale in a Real-Time Marketplace, published by InfoQ.
Pereira said she wanted the audience to leave with two things: how DoorDash built SafeChat, and one architectural pattern usable in almost any AI use case.
DoorDash sees over 4 million messages exchanged in chat every day, over 400,000 calls between consumers and Dashers during deliveries, and over 200,000 images exchanged in chat or SMS.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
First-hand practitioner detail, single unverified source
The engineering account is specific and internally coherent - daily volumes, a 2-to-10-second measured LLM latency range, a sub-100ms/p90 classifier target, an under-10% escalation share, and a documented measurement phase - and it comes directly from the engineer who leads the team. But it is one conference transcript from one publisher with no independent verification, the headline unsafe rate is disclosed only as a band ('small single-digit percent'), no accuracy or false-negative figures are given for the cheap layer, and the successor system is asserted without description.
Live in production at one large operator, no external uptake
SafeChat is disclosed as running in the blocking chat path at over 4 million messages per day, which is real production adoption at meaningful scale, and the preceding instrumentation phase is also described as executed rather than planned. Adoption is nonetheless confined to a single company's self-report: there is no second deployment, no third-party replication of the cascade pattern, no named vendor or tooling, and the system has reportedly already been retired in favour of an undescribed successor.
Engineering detail modest, framing claims run ahead of the record
Mildly overstated overall. The technical substance is unusually restrained - the talk leads with a measured base rate rather than a model, and the numbers given are specific. The overstatement sits in the framing: the discarded system's 'even more powerful' successor is never described in the supplied transcript, and the promise of one pattern usable in 'almost any AI use case' generalises from a single chat pipeline with no reported accuracy or miss-rate data.
Employer-brand conference talk, moderate self-presentation pressure
The account is delivered by a DoorDash engineering lead in a conference presentation about her own team's work, which carries ordinary employer-brand, recruiting, and personal-reputation incentives to present the architecture as a success story - visible in the rewrite framing and the absence of miss-rate data. The publisher's model is to distribute practitioner talks, so editorial adversarialism is low by design. There is no disclosed commercial product, vendor sponsorship, or funding event in the supplied material, which keeps incentive pressure moderate rather than high.
Moderate - specific but single-sourced and incomplete
Confidence is limited by structure rather than plausibility. One publisher, one first-party transcript, no corroboration, key figures given as bands, no effectiveness metrics, and a successor system left undescribed. Against that, the speaker is directly accountable for the system, the numbers are internally consistent (the derived under-10% escalation share matches the stated one), and the derived arithmetic follows from disclosed figures. An internal inconsistency between the supplied ledger and the transcript body further caps confidence in the record's completeness.
product
Also raised $150M for a parts bin, not a bicycle3 distinct publishers
build
Flux moves GitOps' source of truth into registries you own, and mirroring becomes the prerequisite1 distinct publisher
build
JDK 28 firms up: a JSON API in the incubator, and a deprecation notice for Intel Macs1 distinct publisher
product
eSafety tested Roblox's child-safety controls itself, then extracted an enforceable undertaking4 distinct publishers
Distinct publishers with included, body-backed reporting in this cluster.
1 article · August 22, 2026