Science1 publisherNot yet confirmed elsewhere3 min readPublished
Waterloo researchers find chatbots approve existing plans about twice as often as new climate policies
University of Waterloo researchers found 11 chatbots approved existing plans 70% of the time but new climate policies only 34%, across nearly 55,000 prompts. The work measures recommendations to standardized prompts, so how far the lean reaches real purchases and votes is a separate question.
The Scientist · Science desk

What happened
- The prompts came from more than 7,500 queries on car purchases, home heating, diet and municipal or regional climate policy, with each query run on at least six models.
- Models recommended electric vehicles less often than national EV sales shares would suggest, even when the prompt placed the user in Norway.
- The university attributes the lean to models being trained on large historical text collections that reflect past human choices.
Compiled by The ScientistSomething wrong?How this is made
Why it matters
- decision Anyone routing purchase or policy questions through a chatbot now has evidence that labelling an option as new can roughly halve its approval, so prompt and interface wording becomes something to test before launch.
- constraint Pooled rates from unevenly sampled models leave open which model, if any, carries the lean. Choosing between models would take the per-model figures.
- exposure Buyers in markets where the clean option already leads, such as Norway, can receive advice that trails what local buyers already choose.
Seventy over 34 is a ratio of about 2.1, a gap of 36 percentage points [11][12]. According to the University of Waterloo, the policy prompts put the same objective to the models framed either as an existing, active policy or as a new climate intervention or regulatory change [5]. If the wording held everything else constant, the design rules out one alternative explanation: that the models disliked the climate measures themselves. The university says the lean was strongest in political and civic scenarios [10].
Nearly 55,000 responses from more than 7,500 queries works out to about 7.3 models per query [13]. Had every query gone to all 11 models, the total would have been about 82,500 [14]. Each query went to at least six [4]. The pooled 70% and 34% therefore blend models in unequal proportions, and a few strongly biased models could carry much of the gap. The release does not name the 11 models or give rates for each.
The electric vehicle test compares the models with a number measured outside them. Recommendation rates were set against each country's actual EV sales, and the models came in below the market, including when the prompt placed the user in Norway [6]. Sales share describes what buyers already do, so the comparison shows the models lagging the market and leaves open what the best recommendation would have been. The models do adjust somewhat to the user's location, the university says, but their baseline advice still trails real transitions [9].
On cause, the release goes further than the experiment. The experiment measures outputs, so it can show that the lean exists but cannot by itself separate training data from later tuning as the source. The release nonetheless says the bias is "deeply baked into the underlying architecture" of commercial chatbots [16] and links it to training on large historical corpora that reflect past human choices [17].
For a team putting a chatbot in front of customers or officials, the policy result is a framing effect: approval for the same goal fell from 70% to 34% when it was presented as new [2]. A product's system prompt and the way its interface phrases a question sit between the model and the user, and this study shows that wording moves the answer. Seth Wynes, the lead researcher and a professor in Waterloo's Faculty of Environment, pointed users at the same lever. "I think it's worth being aware that these models, even the new ones, have blind spots," he said. "It's good for consumers to be aware of this bias and if you're asking for advice on a topic, you could ask it to make the case for doing something new." [8]
I think the evidence supports a narrower claim than the release's framing that chatbots steer users toward high-emission defaults [15]. The study measured what the models said; measuring how often people follow the advice, and what that does to emissions, would take a different study. In standardized tests, these models favoured existing options and backed a decision the user had already made at double the frequency of alternatives, according to the university [7].
What to watch
- Per-model approval rates and the names of the 11 models, which would show whether the pooled 70% versus 34% gap comes from a few models or all of them.
- The full national EV comparisons, including Norway, to size how far chatbot advice trails actual sales.
- A follow-up that tracks whether people given this advice change purchases or policy positions, or that tests Wynes' suggestion to ask the model to argue for something new.