Skip to content

Leadership1 publisher3 min readPublished

Synthetic consumer panels are typically validated by agreement with human survey answers, not real-world outcomes

BluePill AI's Jeetendra Gangele says validation in the category shows twins reproducing what people say, while evidence that they predict what people buy stays thin. Brands are already querying models in place of six-week, six-figure studies.

The Board Room · Leadership desk

Illustration accompanying Synthetic consumer panels are typically validated by agreement with human survey answers, not real-world outcomes

What happened

  • Jeetendra Gangele, writing for the Forbes Technology Council, says the trouble with AI consumer panels is that fluency gets mistaken for accuracy, and sets out four questions for buyers.
  • A 2025 study of 57 personal care surveys reported synthetic ratings reaching 90% of human test-retest reliability on purchase intent.
  • The Market Research Society's 2024 Delphi report separates statistically synthesized data from LLM-generated participants and warns against assuming the second inherits the first's track record.
  • He points back to Jamieson and Bass, who established in 1989 that stated purchase intent substantially overstates actual trial by margins that vary heavily by category.

Compiled by The Board RoomSomething wrong?How this is made

Why it matters

  • constraint A validation built on panel agreement caps what a six-figure substitution can prove: the buyer gets evidence that twins echo survey answers and keeps the sales risk on its own books.
  • decision The swap has to be scoped by category, because deliberate purchases and impulse buying sit on opposite sides of the method's known weakness.
  • contradiction The two academic citations land on opposite sides of whether models reproduce consumer behaviour, so a business case can cite the literature for or against depending on which paper it picks.
  • exposure The most specific buyer's checklist in circulation comes from a vendor selling one-twin-per-person systems, so a procurement team adopting it wholesale is writing a spec that favours one architecture.

Validation in this category has one standard shape: run the same study with human respondents and with twins, then compare the results [8]. That design is where the 90 per cent figure comes from, and it fixes what the figure can mean. The comparison is against a human panel's agreement with itself on a repeat survey. What is being scored is survey answers.

Gangele, cofounder and CTO of BluePill AI [1], states the limit himself: agreement with a human panel shows that twins reproduce what people say, and not that they predict what people do [9]. The 1989 Jamieson and Bass finding predates the 2025 validation study by 36 years [19]. A synthetic panel that matches human survey answers closely therefore inherits the distance those answers already have from trial. Most of the field, he writes, has considerably more evidence of panel agreement than of performance against real market outcomes [11].

A skeptic would call his four questions a spec sheet for his own product, and the first question supports that reading: he wants systems grounded in long-form interviews with real, consenting people, one digital twin per person, and says a model instructed to act like your target consumer is only being tested on role-play [18]. Two of his other points cut against the seller. Twins built from interviews inherit stated preferences and articulated reasoning, which leaves them strong on deliberate decisions and weak on the impulse buying that drives a large share of packaged goods purchases [14]. A last-second candy purchase at the checkout falls outside what the person can narrate, so it stays out of the interview [15]. "It is an open problem for the whole category, and anyone describing it as solved is likely just trying to sell something," he wrote [16].

The academic record he points to runs both ways. A 2023 paper in Political Analysis found that language models conditioned on demographic profiles could reproduce human survey patterns [4]. A 2024 review in Psychology and Marketing found the same models approximated some consumer responses while failing to reproduce well-documented effects including the endowment effect and mental accounting [5].

For predictive claims he asks for a blind design: concepts predicted without revealing the expected result, then compared with actual sales or trials, with the study design, the metric, the benchmark and the failure cases supplied [13]. "A demo is not a validation," he wrote [12].

The question in front of a research budget this quarter is narrower than whether to keep or retire the human panel. Six weeks and six figures buys a human study, and brands are now querying models instead [2]. On the evidence Gangele cites, the swap holds where the purchase is deliberate and the twin is current, and a twin built before a price shock is a consumer who no longer exists [17]. For impulse categories, he cites no comparison of twin predictions with sales.

What to watch

  • Whether the 57-survey result replicates outside personal care, where purchase intent is comparatively easy to state.
  • Whether the Market Research Society or a buyer-side body turns its 2024 distinction into a disclosure requirement on commissioned studies.
  • Whether any vendor publishes a preregistered sales comparison that includes the concepts its twins scored wrongly.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories