Skip to content

BuildNot yet confirmed elsewhere1 publisher3 min readPublished

Simile's simulated shoppers: two 85 percents, one budget line

A $2B Series B rests on synthetic respondents agreeing with human focus groups 85 to 99 percent of the time. The band's floor and ceiling are not the same product.

The Engineer · Build desk

How we use AISend a correction

Photograph accompanying Simile's simulated shoppers: two 85 percents, one budget line
Photo: latent.space

What happened

  • Founder Joon Sung Park's earlier work built digital twins of 1,000 real people that matched behaviour 85 percent as well as those people matched their own prior answers.
  • The method leans on long-form interviews, transaction and observational data, and post-training on randomised controlled trials.
  • Park's stated ambition runs to simulating all 8 billion people, with worlds large enough to need a data centre.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • cost The money at stake is the human panel line in the research budget, and the buyer paying for synthetic respondents is the same team that would have to fund the human study proving they worked.
  • contradiction One 85 percent is scored against a human panel and the other against people repeating themselves, so a procurement threshold written against the headline figure is measuring something the research...
  • decision Insights leads now have to sort their question types by tolerance for error, because the same vendor number covers work you can act on and work you can only use to narrow a list.
  • capability If behaviour has to be trained into weights rather than prompted, the entrants who can compete are those willing to fund interview and trial collection, not those wrapping a frontier API.

Two numbers in latent.space's write-up look like the same number and are not. The commercial figure is 85 to 99 percent agreement with human focus groups across tens of millions of simulations for Fortune 100 buyers [10]. The research figure, from Park's digital twins of 1,000 real people, is 85 percent as accurate as those people were at reproducing their own earlier answers [1]. The second is scored against human self-consistency, which is not a ceiling of 100. Anyone reading the first number as a straight hit rate against ground truth is reading a different measurement than the one the underlying work reports.

Then the band itself. At 85 percent agreement, roughly one answer in seven diverges from the human panel; at 99 percent, one in a hundred. The disagreement rate at the bottom of the quoted range is about fifteen times the rate at the top [9]. That is not a tolerance, it is two different products sold under one figure. A concept screen that only needs rank order survives one in seven; a claim test that decides packaging copy or a price point does not. The source does not say which question types land where in the band, and that mapping is the whole procurement question.

The mechanism is more interesting than the round. Park's position is that prompting frontier models does not get you there, and that capturing what he calls social physics may require changing model weights [3]. Models optimised to behave rationally make poor simulations of people who do not [11]. The inputs are long-form interviews, observational and transaction data, and post-training on randomised controlled trials [4], on the argument that web text records what people say rather than what they do [12]. If that is right, the defensible asset is the collection apparatus and the post-training run, not the inference call. A competitor with an API key and a persona prompt is building a different thing, and cheaply.

That also sets the honest limit. The interview lists evaluating simulations, rather than stacking model hallucinations, as an open problem in its own right [5]. A vendor whose accuracy claim is expressed against human panels needs human panels to keep proving the claim, which means the cheapest version of this business still buys some of the thing it replaces. Park frames market research as the starting point rather than the market [6], which is the usual signal that the near-term revenue and the long-term story are not the same slide.

One more thing worth naming: the write-up calls it a $2B Series B backed by GreenOaks and Index Ventures, with Fei-Fei Li and Andrej Karpathy among the backers [2], and does not say whether $2B is money in or the valuation attached to it [8]. Those are very different companies. Until that resolves, the number to hold onto is the fifteen-fold spread inside the accuracy claim, not the one with the B on the end.

What to watch

  • Whether the $2B figure resolves as capital raised or post-money valuation, which changes how much runway the collection apparatus actually has.
  • A published breakdown of where in the 85 to 99 percent band each question type falls, ideally with the human panels held back as a scoring set.
  • Whether any Fortune 100 buyer names a category where it has retired human panels outright rather than running them in parallel.

Clarity's read

What the record supports and how the coverage leans. The claims behind it follow.

Reality

Evidence24
Adoption28
Hype gap+46
Incentives78
Confidence52
Why these scores

Claim ledger

Ranked by verification strength, evidence, and original report placement.

  1. [1]

    Park's research created digital twins of 1,000 real people that reproduced human behaviour and attitudes 85% as accurately as those people reproduced their own responses.

    ReportedSupportedSource: Joon Sung Park via latent.space2 sources— create a free account to open themView cited source
  2. [2]

    Simile AI has a $2B Series B backed by GreenOaks and Index Ventures, with backers including Fei-Fei Li and Andrej Karpathy.

    ReportedSupportedSource: latent.spaceView cited source
  3. [3]

    Park argues that understanding social physics may require changing model weights rather than simply prompting frontier LLMs.

    ReportedSupportedSource: Joon Sung ParkView cited source

Sources

1 independent publisher whose own reporting we read for this story.

  1. latent.space

    1 article · August 21, 2026

    Simulation: the new Scaling Law — Joon Sung Park, Simile AI

Share your take

Let Clarity write the post for you.

Signed-in readers get a short post drafted on this story in the register they choose — narrative, analytical, or a direct position — editable to the last word before it goes anywhere. The share buttons at the top of this story work without an account.

Topics and entities

Follow any of these and your For You feed starts watching them — no settings page required.

Topics

Entities

Loading related stories