Science1 distinct publisher2 min readPublished
LLM surrogates of more than 2,000 real respondents were wrong about a quarter of the time and landed roughly where a chatbot given demographics alone landed, with errors that pull toward stereotype.
The Scientist · Science desk

Compiled by The ScientistSomething wrong?How this is made
Errors with a direction are the kind that survive aggregation. Noise washes out when you average over more synthetic respondents; a pull toward the modal answer does not. That is what the paper documents: twin responses came out more homogeneous than the responses of the people they stood in for, and they skewed toward demographic stereotypes [7].
The elicitation ledger is worth spelling out. Roughly 500 questions each for something over 2,000 people is on the order of a million answers, and both figures are stated as minimums, so that is a floor [19]. Many of those questions came from scales that psychology, economics and business research already lean on to characterize a person, and respondents also sat online tests of thought patterns and biases [11]. Toubia says the resulting dataset has been downloaded about 25,000 times [13].
Against that input, the twins were wrong on average about a quarter of the time [4]. Where the profiles did earn something was spread: given two people who rate their self-control 2 and 4, a model holding only demographics tends to answer 3 for both, while the twin answers 3 and 5, wrong in both cases but not flattened [6]. Olivier Toubia, a computational social scientist at Columbia Business School, allows that "there's some promise" while calling the twins' performance "overall a bit disappointing" [12]. For anyone estimating an effect size, recovered variance without accuracy is the more treacherous of the two failures, because a synthetic sample can reproduce the shape of a distribution while placing the wrong people in its tails. That inference is mine, not a result in the paper.
The thing a quarter-wrong average does not tell you is which quarter. It pools 19 experiments spanning questions as different as how people judge donors who give to both Republican and Democratic candidates and what they think about algorithmic hiring [3]. A substitution program does not run into the average hit rate; it runs into whether the surrogate errs the same way inside the one subgroup the study is about.
Hadi Hosseini of Penn State, who has watched LLM agents make health care decisions under scarcity, reports the same directional pull in his own domain: the models distort judgment toward something very rational and more reasonable than people actually are [14]. Two labs finding the same skew in unrelated tasks is the part of this that generalizes. Toubia's own caution points at his field rather than at the models. Social scientists tend to assume their scales capture the full range of human experience, he says, and predicting human behavior with a machine is incredibly hard [18].
Ranked by verification strength, evidence, and original report placement.
A study appearing September 2 in Science Advances suggests AI twins designed to mimic a given individual's behavior instead distort their surrogate's views, creating a "funhouse mirror" effect, the researchers note.
To develop the twins, the researchers fed all the information for each individual into a large language model and prompted the LLM to respond as if it were that person.
Across 19 social science experiments, the team evaluated everything from how individuals and their twins respond to people who donate to both Republican and Democratic party candidates to what they "think" about algorithmic hiring.
The digital twins performed better than chance, the team found, but they were wrong on average about a quarter of the time.
The twins performed roughly on par with chatbots that received demographic information alone.
The digital twins better captured real variation in people's responses than LLMs that knew only demographic info: where one person rates self-control a 2 and another a 4, the limited-info LLM might report 3 for each, washing out differences, while the twin might report 3 and 5, still wrong but giving a better sense of potential differences across the group.
Distinct publishers with included, body-backed reporting in this cluster.
1 article · September 3, 2026
Follow any of these and your For You feed starts watching them — no settings page required.
product
Penn State imaged 100 meters of subsurface using lightning and buried telecom fiber2 distinct publishers
science
Corals hosting heat-tolerant algae lost more tissue when infection followed the heat1 distinct publisher
science
Sheeppox DNA pulled from the York Gospels turns library parchment into a pathogen archive2 distinct publishers
science
Comparative brain data recasts evolution as a tug of war between two wiring styles1 distinct publisher
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
One peer-reviewed paper, retold by one desk
Every figure here — 19 experiments, 2,000-plus respondents, a quarter wrong — reaches readers through Science News' reading of a single Science Advances paper. What lifts it above ordinary single-source retelling is that the deflationary quotes come from the study's own author and that an unaffiliated Penn State researcher reports the same rationality skew in unrelated work on scarce health care resources. What is missing is grain: no intervals, no spread across the 19 experiments, and not even the name of the model doing the impersonating.
Downloads of a questionnaire, not twins in service
The lone traction figure belongs to the input, not the product: roughly 25,000 pulls of the free profile dataset, offered as "like 25,000 times" by the man who published it. Nobody in this reporting runs synthetic respondents in anger. The two uses floated — long-form answers from tireless twins, and pretesting a design before burning human patience — are proposals from the study's author, and the substitution driving the whole field is described as sitting on wish lists.
The framing runs cooler than the finding
This is the unusual case where the write-up sells the result short. "Wrong about a quarter of the time" reads like a tuning problem you fix next quarter. The mechanism Science News actually describes is harder than that: answers compress toward demographic stereotype, and accuracy climbs with the participant's income and education — a bias result, tucked mid-paragraph, that outlives whatever model was used.
The best-quoted skeptic built the thing
Toubia authored the study, assembled the 500-question dataset it runs on, volunteers its download count, and lands on wanting "more complex training methods" plus two jobs his current twins could still do — a critique shaped, quite openly, like a research agenda. Hosseini corroborates from his own parallel project and then prescribes the remedy that would extend it: chatbots shadowing people through the day. All of this is disclosed rather than hidden, and nobody with revenue riding on synthetic panels is quoted at all.
Narrow claims, named sources, no second look
One outlet, one paper, two researchers, and no independent verification of the underlying numbers keeps this well short of settled. It holds up better than that arithmetic suggests because the claims are specific and attributed by name, and because the load falls on statements that work against the interests of the people making them. Treat the direction as solid and the exact quarter-of-the-time figure as provisional until the paper is read on its own terms.