Build1 distinct publisher3 min readUpdated
A $2B Series B rests on synthetic respondents agreeing with human focus groups 85 to 99 percent of the time. The band's floor and ceiling are not the same product.
The Engineer · Build desk

Compiled by The EngineerSomething wrong?How this is made
Two numbers in latent.space's write-up look like the same number and are not. The commercial figure is 85 to 99 percent agreement with human focus groups across tens of millions of simulations for Fortune 100 buyers [2]. The research figure, from Park's digital twins of 1,000 real people, is 85 percent as accurate as those people were at reproducing their own earlier answers [4]. The second is scored against human self-consistency, which is not a ceiling of 100. Anyone reading the first number as a straight hit rate against ground truth is reading a different measurement than the one the underlying work reports.
Then the band itself. At 85 percent agreement, roughly one answer in seven diverges from the human panel; at 99 percent, one in a hundred. The disagreement rate at the bottom of the quoted range is about fifteen times the rate at the top [3]. That is not a tolerance, it is two different products sold under one figure. A concept screen that only needs rank order survives one in seven; a claim test that decides packaging copy or a price point does not. The source does not say which question types land where in the band, and that mapping is the whole procurement question.
The mechanism is more interesting than the round. Park's position is that prompting frontier models does not get you there, and that capturing what he calls social physics may require changing model weights [5]. Models optimised to behave rationally make poor simulations of people who do not [6]. The inputs are long-form interviews, observational and transaction data, and post-training on randomised controlled trials [7], on the argument that web text records what people say rather than what they do [8]. If that is right, the defensible asset is the collection apparatus and the post-training run, not the inference call. A competitor with an API key and a persona prompt is building a different thing, and cheaply.
That also sets the honest limit. The interview lists evaluating simulations, rather than stacking model hallucinations, as an open problem in its own right [9]. A vendor whose accuracy claim is expressed against human panels needs human panels to keep proving the claim, which means the cheapest version of this business still buys some of the thing it replaces. Park frames market research as the starting point rather than the market [10], which is the usual signal that the near-term revenue and the long-term story are not the same slide.
One more thing worth naming: the write-up calls it a $2B Series B backed by GreenOaks and Index Ventures, with Fei-Fei Li and Andrej Karpathy among the backers [1], and does not say whether $2B is money in or the valuation attached to it [13]. Those are very different companies. Until that resolves, the number to hold onto is the fifteen-fold spread inside the accuracy claim, not the one with the B on the end.
Follow any of these and your For You feed starts watching them — no settings page required.
Ranked by verification strength, evidence, and original report placement.
Park's research created digital twins of 1,000 real people that reproduced human behaviour and attitudes 85% as accurately as those people reproduced their own responses.
Simile AI has a $2B Series B backed by GreenOaks and Index Ventures, with backers including Fei-Fei Li and Andrej Karpathy.
Park argues that understanding social physics may require changing model weights rather than simply prompting frontier LLMs.
Simile's approach uses long-form interviews, observational and transaction data, randomised controlled trials, population-level and individual-level models, and post-training on the causal mechanisms behind decisions.
The interview treats how to evaluate simulations, rather than simply stacking LLM hallucinations, as a distinct open problem.
Park frames market research as only the starting point for simulation.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Single vendor-facing interview, no methodology
All facts trace to one publisher's founder interview. The commercial accuracy figure has no stated methodology or independent check, the prior research paper is not supplied, and the source itself frames simulation evaluation as an open problem. Only the internal arithmetic on the quoted band and the description of the method stack are verifiable from the text.
One named client, self-disclosed volume
There is real disclosed usage - tens of millions of simulations and Fortune 100 clients - but only CVS is named, the volume is self-reported, and no procurement, pricing, or renewal detail appears. That is early commercial traction as described by the vendor, not independently observed adoption.
Ambition priced well ahead of measurement
The framing runs from a 14-point accuracy band to simulating all 8 billion people and data-centre-scale worlds, while the source simultaneously admits that evaluating simulations is unsolved and never says where inside the band client work lands. The gap is overstatement of demonstrated capability rather than fabrication: the disclosed research result and method stack are concrete, but the headline range and the civilisation-scale ambition outrun what is measured here.
Founder promotion around a funding moment
The only source is a founder appearing on an access-driven AI podcast at the moment a large round is announced, with hiring and company links appended. Both parties benefit from the category being described as a 'Second Summer' of simulation, and every commercial number originates with the seller.
Text is clear, corroboration is absent
What the source says is unambiguous and directly quotable, so claims about the narrative and the arithmetic on the band are solid. Confidence is capped in the middle because a single self-interested publisher supplies every external fact, so the underlying commercial and technical realities cannot be triangulated.
invest
Three rounds, $2.31B, and one experiment where blocking 1% of Manhattan broke the model1 distinct publisher
leadership
The year's most useful AI tool at one desk was a folder full of Markdown1 distinct publisher
build
The $559M-versus-$12.3B quarter matters more than the $65B run rate4 distinct publishers
product
Also raised $150M for a parts bin, not a bicycle3 distinct publishers
Distinct publishers with included, body-backed reporting in this cluster.
1 article · August 21, 2026