Published Invest3 min read
Your model's politics are a language setting you cannot inspect
A Nature study finds Chinese state media in open training data and measurably different answers in Chinese; Meta's Oversight Board finds four labs' models refuse to criticize repressive governments at twice the rate.
Context for builders, not their beat.See today for builders

What happened
- A peer-reviewed study published in Nature found evidence that Chinese state-controlled media makes its way into AI training data and can influence how models answer questions about China.
- Meta's independent Oversight Board found models from Anthropic, OpenAI, Google and Meta were more than twice as likely to refuse requests to criticize governments in countries that restrict political speech than in freer countries.
- The Nature researchers identified more than three million Chinese-language documents in the open-source training dataset CulturaX, which is used to train and improve LLMs.
- Claude Sonnet, Claude Opus, GPT-3.5 Instruct, GPT-4 and GPT-4o could reproduce distinctive phrases from Chinese state-coordinated media at rates ranging from 3% to nearly 10%, which researchers took as a sign the models had encountered the material during training.
- After further training Meta's open-weight Llama 2 13B, picked because it had very little to zero Chinese state media in its training data, on just 6,400 Chinese state-scripted news examples, the model produced a more Beijing-friendly answer than the baseline model nearly 80% of the time.
Compiled by The InvestorSomething wrong?How this is made
Why it matters
Two pieces of research arrived at the same defect from opposite ends: a peer-reviewed Nature paper found evidence that Chinese state-controlled media enters AI training data and shifts how models answer questions about China [1], while Meta's independent Oversight Board found that models from Anthropic, OpenAI, Google and Meta were more than twice as likely to refuse a request to criticize a government in countries that restrict political speech than in freer ones [2]. For anyone procuring these systems, "politically neutral" turns out to describe a behaviour that varies with the language of the prompt [8] and with the press-freedom score of the country being asked about [9].
The training-data half is the more concrete. Researchers identified more than three million Chinese-language documents in CulturaX, an open-source dataset used to train and improve LLMs [3]. Claude Sonnet, Claude Opus, GPT-3.5 Instruct, GPT-4 and GPT-4o could reproduce distinctive phrases from Chinese state-coordinated media at rates from 3% to nearly 10%, which the authors read as evidence the material was encountered in training [4]. They then established causation on a model they could actually touch: Llama 2 13B, chosen because it had very little to zero Chinese state media in its data, produced a more Beijing-friendly answer than the baseline nearly 80% of the time after further training on just 6,400 state-scripted news examples [5]. At ten times that dose, asked whether China is an autocracy, the baseline said yes and the retrained model described the country as democratic, invoking the Communist Party's concept of "people's democracy" [6][16].
They could not run that experiment on OpenAI's or Anthropic's systems, whose training is largely opaque, so they asked the same political questions in both languages instead [7]. The Chinese-language answer was rated more favorable to Chinese leaders and institutions 68.8% of the time for Claude Sonnet, 72.6% for GPT-3.5, 84% for GPT-4o and 88.2% for Claude Opus [8]. Every model tested tilted in the majority of comparisons, with a 19.4-point spread between the two Claude versions [17]. Across 6,051 prompts covering 37 countries, lower press freedom correlated with more favorable descriptions when GPT-3.5, GPT-4o, Claude Opus and Claude Sonnet were queried in the local dominant language rather than English [9].
The Oversight Board tested 10 commercial models from Anthropic, DeepSeek, Google, Meta, OpenAI and xAI with identical prompts about five restrictive-speech countries including China, Saudi Arabia, Thailand, Turkey and Cambodia, against five permissive ones including the United States, the United Kingdom, Japan and Taiwan [13]. It calls the result "censorship-by-proxy": models sometimes acting as though one country's political restrictions applied to users outside it [14].
Neither study shows Beijing deliberately manipulated any of the four American labs [11]. That is the uncomfortable part. The Nature authors describe the mechanism as severing information from its source, "effectively laundering government-manipulated content into ostensibly objective text" [10]. No adversary action is required for a scraped corpus to carry a state's framing into a product sold as even-handed. Anthropic has reported a 94% score on its own political even-handedness evaluation [12]; that evaluation was not the one that found an 88.2% tilt in Claude Opus [8].
Claim ledger
Ranked by verification strength, evidence, and original report placement.
- [1]
A peer-reviewed study published in Nature found evidence that Chinese state-controlled media makes its way into AI training data and can influence how models answer questions about China.
- [2]
Meta's independent Oversight Board found models from Anthropic, OpenAI, Google and Meta were more than twice as likely to refuse requests to criticize governments in countries that restrict political speech than in freer countries.
- [3]
The Nature researchers identified more than three million Chinese-language documents in the open-source training dataset CulturaX, which is used to train and improve LLMs.
ReportedView cited source - [4]
Claude Sonnet, Claude Opus, GPT-3.5 Instruct, GPT-4 and GPT-4o could reproduce distinctive phrases from Chinese state-coordinated media at rates ranging from 3% to nearly 10%, which researchers took as a sign the models had encountered the material during training.
ReportedView cited source - [5]
After further training Meta's open-weight Llama 2 13B, picked because it had very little to zero Chinese state media in its training data, on just 6,400 Chinese state-scripted news examples, the model produced a more Beijing-friendly answer than the baseline model nearly 80% of the time.
ReportedView cited source - [6]
After 64,000 state-scripted examples, the retrained model was asked whether China is an autocracy; the baseline model said it was, while the state-scripted version described China as democratic and invoked the Chinese Communist Party's concept of "people's democracy."
ReportedView cited source
Sources & coverage · 1 publisher
The reporting this story was synthesized from, earliest first. Every link goes to the original.
- fortune.comMia OsmonbekovAug 13‘Multi-part case study on China’s media’ finds that AI models can’t hallucinate away Chinese censorship
Additional citations
- Nature study, via Fortune
- Meta Oversight Board report, via Fortune
- Nature researchers, quoted by Fortune
- Anthropic, via Fortune


