Skip to content

Build1 publisher3 min readPublished

DoorDash threw away a working safety classifier. The measurement that came first is the reusable part

Bruna Pereira's SafeChat talk puts a number on unsafe chat traffic before it puts a model in the message path. That ordering, not the model, is what survives a rewrite.

The Engineer · Build desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

Photograph accompanying DoorDash threw away a working safety classifier. The measurement that came first is the reusable part
Photo: infoq.com

What happened

  • A DoorDash engineer told an InfoQ audience the team built a chat safety system, got it working in production, and then threw it away for a more powerful replacement.
  • The platform carries over 4 million chat messages a day, plus over 400,000 consumer-to-Dasher calls during deliveries and over 200,000 images.
  • Chat allows a fraction of a second to rule a message safe; the team measured LLM calls averaging 2 to 10 seconds, with wide variance.
  • Before building anything, the team spent a couple of months instrumenting chat and running a free moderation API against it asynchronously.
  • That measurement found only a small single-digit percent of messages were unsafe, and the classifier was trained on the data it produced.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • constraint Accuracy cannot buy back time, so anything that answers in seconds is disqualified from the message path outright. That makes this a routing problem, not a model bake-off.
  • decision The base rate on your own traffic is the gate before the build. Without that number, preferring an aggressive cheap filter over paying per message is a hunch you cannot defend in review.
  • capability Because the escalated share is tiny, the slow expensive model becomes affordable where it matters: the residue needs a couple of dozen concurrent calls, not thousands.
  • exposure A layer that is right almost all of the time pushes its residual failures onto the people the system exists to protect, and the talk publishes no false-negative rate for it.

Nine percent is a generous ceiling on "small single digit", and it is enough to do the architecture arithmetic. Of 4 million daily messages, at most 360,000 would ever need an expensive opinion, and at least 3.64 million would not [15]. At the 2 to 10 second latencies Pereira quotes for her team's LLM calls [9], that escalated slice works out to roughly 8 to 42 model calls in flight at any moment, averaged across the day [16]. Peaks will be worse than the average, but that is a fleet you can reason about, and it is why the cheap layer is not a compromise: it turns an unaffordable per-message problem into a small one you can afford to do well.

The latency arithmetic is less forgiving. A sub-second decision budget against 2 to 10 seconds of measured model latency is over budget by 2x at best and 10x at worst [17], and no gain in classification quality returns any of that time. Model selection is not the lever here. Routing is.

Which is what makes the rewrite the more useful half of the talk. The published transcript stops as the small classifier's three jobs are introduced and never lists them [18], so the replacement's internals are not in this record. What Pereira says she wants the audience to keep is a pattern rather than an implementation, one she claims applies to almost any AI use case [5]. On the evidence she has already laid out, that pattern is an ordering: get the base rate off your own traffic first, then size a cheap filter to it and let the expensive model see only what the filter cannot settle. None of that is bound to a particular model, which is exactly why a system that worked could be thrown out [1] without taking the design with it.

The bill for the pattern lands in one place. "Right almost all of the time" [3] on a safety product means the remainder is an abusive message delivered, to two parties whose relationship runs somewhere between six and forty minutes and who have no history to fall back on [13]. The talk gives no false-negative rate for the cheap layer. Against a stated goal that every message reaching its destination has been classified as safe [7], that is the figure worth asking for, and it is a decision about tolerable residual risk rather than a modelling detail. Verbal abuse is described as a meaningful share of incidents on the platform, and feeling safe is treated as a product metric on equal footing with being safe [14], which means the misses are measured in both.

What to watch

  • What replaced SafeChat, and whether the escalation share held once the second system was carrying live traffic.
  • Whether voice calls and images get the same cheap-filter-first treatment, or need separately measured base rates and latency budgets.
  • Any published false-negative rate for the cheap layer, which is the number that decides whether an aggressive filter is defensible on a safety product.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories