Product1 distinct publisher3 min readPublished
Health is among the most common things people ask chatbots. The teams that define what counts as a safety failure sit in a handful of Western offices, and their evaluations never test the dialect a user types.
The Product Desk · Product desk

Compiled by The Product DeskSomething wrong?How this is made
A patient reading that they have been given intravenous insecticides has no way to tell whether the mistake came from the clinic or the translation layer. In Tigrinya, spoken by about 9 million people in Eritrea and northern Ethiopia [17], machine translation produced that sentence in place of "you have been given intravenous antibiotics," turned smallpox into syphilis and gonorrhea into diabetes [18]. Elizabeth Orembo, a fellow at Research ICT Africa, told Rest of World that mistranslations like these can be life-threatening [19]. Researchers examining natural language processing in healthcare in Africa found linguistic and cultural bias alongside poor adaptation to medical contexts and outright translation errors [16].
Orembo's account of why that survives review is the part worth borrowing. Large AI firms concentrate on model risks: deception, autonomous behavior, cyber capabilities, aiding bioweapons [12]. Deployment risks get much less attention, among them discrimination, exclusion, surveillance, language failures, and the inability of affected communities to get anything remediated [13]. A risk taxonomy is also a test plan. What sits outside it will not be found by the people paid to look.
New-market teams often assume that a multilingual model has the language covered, but users write the way they speak and carry the urgency in the phrasing, not in the vocabulary a benchmark tests. A review in India found that more than two-thirds of chatbots did not adequately account for dialects or recognize urgency cues [6], which leaves under a third that did [7]. That is a depth-of-use number rather than a volume number, and it is the metric that belongs in a launch review, measuring how often the assistant missed the signal that a question was an emergency rather than how many sessions the pilot generated.
The second failure mode is environmental. Trust and safety evaluations assume reliable electricity and internet connectivity, functioning courts, robust data protection laws, formal labor markets, and a press and civil society that report failures [14]. Where several of those are absent, the support thread that would tell you the product is hurting someone never gets written, which is how, in Orembo's words, "a model can pass every frontier safety evaluation and still produce unsafe outcomes when deployed" [15]. The UN has reported that risks fall disproportionately on developing nations, which adopt AI more slowly but have inadequate resources, limited domestic AI infrastructure and more dependence on foreign technology [10].
There is no universal trust and safety standard to procure against - each company writes its own framework and its own enforcement methods [9], so procurement alone cannot settle this. Urvashi Aneja of Digital Futures Lab, who is preparing a report for the UN on AI safety in developing nations, told Rest of World that the evaluation frameworks, the standards and the overseeing institutions have largely been designed in and for a small set of high-income countries [8].
So this is for the person shipping a health-adjacent assistant into six locales next quarter and answering for it when a clinic complains. Two questions per locale, each answered with a name rather than a percentage. Has someone tested the language as users actually write it, dialect and urgency phrasing included, on questions where a wrong answer changes a treatment decision? And when it fails, does the user reach a human, and does the complaint land somewhere a person reads it? A locale that cannot answer either question with a name is unsupported, whatever the launch tier calls it, and the localization work that would change the answer sits in the safety budget rather than the growth one.
Ranked by verification strength, evidence, and original report placement.
Generative AI systems are failing at basic tasks outside Western nations because trust and safety teams are largely concentrated in Silicon Valley and do not reflect the concerns of countries with different languages and cultural contexts.
Urvashi Aneja, founder of Digital Futures Lab and author of a forthcoming UN report on AI safety in developing nations, told Rest of World that the global majority still remains at the margins of AI safety discourse, and that the frameworks used to evaluate AI systems, the standards that govern them and the institutions that oversee them have largely been designed in, and for, a small set of high-income countries.
There is no universal trust and safety standard; each company has its own framework, based on its values and principles, and its own enforcement methods.
The UN said in a recent report that developing nations are adopting AI more slowly than wealthier nations, yet risks fall disproportionately on them because of inadequate resources, limited domestic AI infrastructure and dependence on foreign technologies.
Elizabeth Orembo of Research ICT Africa said big tech firms tend to focus on model risks such as deception, autonomous behavior, cyber capabilities and aiding bioweapons.
Orembo said big tech firms do not pay much attention to deployment risks including discrimination, exclusion, surveillance, language failures, and the inability of affected communities to seek remediation.
Distinct publishers with included, body-backed reporting in this cluster.
1 article · September 1, 2026
Follow any of these and your For You feed starts watching them — no settings page required.
product
Washington's secret AI test is coming for open weights, and release dates go with it2 distinct publishers
invest
Meta's $17 billion settlement puts a cash number on the tech backlash1 distinct publisher
invest
ChatGPT ads reach 40% of this year's revenue target in their first 200 days7 distinct publishers
leadership
Meta scraps Project OT, underscoring how hard AI deployment really is1 distinct publisher
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Named institutions hold the middle; the ends float
The argument's spine is well sourced — Urvashi Aneja on the record, Orembo at Research ICT Africa, Thakur at George Washington University, plus UN, UNDP, and Future of Life Institute documents. The two most quotable numbers are not: 'more than two-thirds of chatbots' comes from a review in India with no name attached, and the Tigrinya mistranslations come from 'researchers' with no study cited. Add an opening that asserts a first-ever training pause and models loose on the internet with nothing but a post on X behind it, and half the memorable material in this story cannot be followed anywhere.
Harms illustrated, never counted
What real-world traction exists here is categorical: health questions are 'among the most common' chatbot uses, ID systems 'lead to' denial of wages and meals, tools misidentify crops. Only two numbers appear in the whole story — two-thirds of chatbots and nine million Tigrinya speakers — and neither tells you how many people are actually routing a symptom description through a model in a language it was never evaluated in. The UN's own finding that these countries adopt AI more slowly cuts against the scale implied by the framing, and the story does not reconcile the two.
Frontier drama oversold, clinic-level harm undersold
Two directions cancel unevenly. The top of the story sells escaped models and an unprecedented pause on the strength of one social post — more weight than the sourcing carries. The bottom, where a translation engine turns antibiotics into insecticides for nine million Tigrinya speakers, is stated flatly and then left, when it is the finding a reader will remember. Net slightly overstated, but the overstatement is in the packaging, not in the substance.
Advocacy on the record, companies unasked
Everyone quoted has a stake and Rest of World says so: Aneja is writing the UN report the piece also leans on as authority, Orembo works at a think tank built to press this case, Thakur runs an initiative named for the argument, and the Future of Life Institute both publishes the scorecard and campaigns on the risks it scores. That is disclosed interest, not hidden interest. The asymmetry is on the other side — Anthropic, Meta, DeepSeek, xAI, and Mistral are graded here, and none were asked to answer.
One byline, dense attribution, one soft opening
Confidence sits in the middle for a reason: the mechanism the story describes — evaluations built on assumptions about courts, power, and press that do not hold where the tools are used — is coherently argued and attributed to people who would know. But nothing in it has been checked by a second newsroom, and the passage most likely to be repeated elsewhere is the one with the least behind it.