Skip to content

Science1 publisher2 min readPublished

A Nature analysis sorts LLM human proxies into four roles that each need their own validity test

The analysis argues that a model standing in for people can be sound for one research purpose and unsound for the next, and that the four purposes it identifies each rest on a different claim about human likeness.

The Scientist · Science desk

Photograph accompanying A Nature analysis sorts LLM human proxies into four roles that each need their own validity test
Photo: nature.com

What happened

  • An analysis published on nature.com organizes research that uses large language models as human proxies into four roles: believable agents, task agents, experimental subjects and silicon samples.
  • Comparing the roles against each other is the method, and the conclusion the authors draw is that similarity to humans is not a single feature.
  • Among the work being classified is Park and colleagues' 2023 study of generative agents as interactive simulacra of human behavior, presented at the ACM user interface software symposium.
  • Only the abstract sits outside the paywall; nature.com prices the article at $39.95 and a year of the journal at $119 for 12 digital issues.

Compiled by The ScientistSomething wrong?How this is made

Why it matters

  • constraint A single human-likeness score cannot certify a proxy across all four uses, so a team reporting one number has left the other three claims unmeasured.
  • decision Anyone buying synthetic respondents has to fix which of the four claims they are making before they can say what would count as validation. That choice picks the metric.
  • precedent In peer review, saying a proxy was validated against humans stops being a complete answer, because a reviewer can now ask which role and which criterion.

Four roles imply four measurements. Passing one says little about the others. A believable agent is judged by an observer's sense that the behavior looks like a person's. That is the tradition of Bates's 1994 paper on the role of emotion in believable agents and Loyall's 1997 Carnegie Mellon thesis on building interactive personalities, both cited [10][11]. A task agent is judged on whether the work got done, the framing of the multi-agent systems and AI textbooks also in the list [12]. Both of those tests can pass without telling you whether a model's answers track a population.

The other two roles do. Argyle and colleagues' 2023 paper in Political Analysis asked whether a language model can simulate human samples, which is a distributional claim about matching a target population [5]. The experimental-subject role asks a causal question, and Horton, Filippas and Manning's preprint puts it as what can be learned from "homo silicus" [9]. Those two come apart. A model can reproduce the marginals of a survey that sat in its training data and still shift the wrong way when the wording of a treatment changes.

The criteria carry more weight here than in earlier simulation work, because of the abstract's second sentence. LLMs do not explain the mechanisms of human behavior, and the case for using them rests on their standing in for people [1]. Older traditions did claim mechanism, and several are cited, including Anderson and Lebiere's 1998 book on the atomic components of thought and Laird's book on the Soar cognitive architecture [14]. So are Schelling's 1969 segregation models and Epstein and Axtell's Growing Artificial Societies [13][20]. The cited work spans 57 years, from Schelling to a preprint dated 2026 [18]. Tversky's 1977 "Features of similarity" is in the list too [17].

The public abstract names the roles and states that each requires its own criteria; it does not print which criterion belongs to which role [22], or estimate how much of the existing proxy literature is checked against the wrong one [23]. The nearest thing to that in the record is one of its own references: Gao, Lee, Burtch and Fazelpour's 2025 paper in PNAS titled "Take caution in using LLMs as human surrogates" [7]. Another is Dillion and colleagues' 2023 question in Trends in Cognitive Sciences about whether AI language models can replace human participants [8].

None of the four criteria is about cost. Whether a synthetic panel is cheaper than fielding a real survey is a separate question from whether its answers are valid. A team can get a clean answer on validity while the price per generated respondent decides the project.

What to watch

  • Whether the full text assigns a specific criterion to each of the four roles, since the abstract only asserts that the criteria differ.
  • Whether vendors of synthetic-respondent panels start reporting out-of-sample distributional match separately from treatment-effect direction.
  • Whether journals begin asking authors which of the four claims a paper is making before assessing its validation.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories