Invest1 publisher3 min readPublished
Garry Tan relocates AI's moat from the model weights to the price list
Y Combinator's chief executive would leave distillation alone and have regulators police the gap between open weight and frontier pricing, while the NSA, CISA and FBI treat the copying as a cyber security matter.
The Investor · Invest desk

What happened
- Garry Tan said "I would do nothing" about distillation, telling CNBC's Kate Rooney at Y Combinator's Demo Day that there is instead a case for an American distillation regime.
- The NSA, CISA and FBI released an official cyber security advisory on distillation on Tuesday, in the middle of the frontier labs' campaign against Chinese copying of their models.
- Tan pointed at the copyright problem sitting under the labs' complaint, since much of the data used to train frontier models may itself be covered by copyright law.
Compiled by The InvestorSomething wrong?How this is made
Why it matters
- constraint Tan's condition sets a corridor for frontier pricing: the premium has to stay wide enough to fund training runs and narrow enough that open weights do not take the volume.
- exposure A lab seeking protection for its outputs is a defendant on its inputs, so a copyright ruling on training data reaches the same weights it wants defended.
- decision With roughly three-quarters of the Demo Day cohort in AI, Y Combinator's portfolio economics run through inference prices, and any regime that props up frontier pricing raises its companies' cost of goods.
- contradiction Washington is treating distillation as a security problem while Tan treats it as a pricing problem, and the published account of the advisory settles neither because it does not say what the agencies want done.
Tan's condition is where the money is. CNBC reports that he wants regulators to build an equilibrium between open weight models and frontier models, on the condition that frontier models keep a price premium big enough to make their business model feasible [10]. That puts the defensible asset in the price list, or rather in the size of the gap between two price lists. Distillation, as the practice is defined, uses the outputs of a more capable model to train a smaller or less capable one [3]. A rule against it is a rule protecting that gap.
Then there is Y Combinator's own book. Of the 196 startups presenting at Demo Day, 149 were categorised as machine learning and AI ventures [18], about 76 per cent of the cohort, with 47 outside it [19][20]. Those 149 are buyers of inference. Open weights that come within reach of frontier capability cut their cost of goods, and Tan runs the firm that owns a slice of each of them.
The copyright point is the one doing real work, and Tan raised it: much of the data used to train these models may be covered by copyright [7]. The New York Times sued OpenAI and Microsoft in 2023 over unauthorised use of its articles as training data [8]. A consortium of book authors settled with Anthropic in 2025 on a similar claim [9]. A lab asking Washington to police what comes out of its models is simultaneously defending how it got what went in.
What this record does not contain is the advisory. The NSA, CISA and FBI released an official cyber security advisory on distillation on Tuesday [6], and CNBC's account does not say what it recommends or whether it names a company [21]. OpenAI's side of the case is stated as a belief, that DeepSeek's V3 and R1 architectures were distilled from GPT-4 and GPT-4o [5]. Anthropic has named Moonshot AI, DeepSeek and MiniMax [4].
On safety Tan drew the line at operational risk. "We need to be focused on science fact, not science fiction," he said [13]. "We need to be responding to what is happening right now. If there was a breach and a coordinated attempt by agents to take over our infrastructure, what do we do about it?" [14] Anthropic said on Thursday that it had blocked Claude access in five cases of scientists in unspecified foreign countries researching dangerous pathogens, out of concern they were covertly building bioweapons [16].
There are three ways this goes from here. Enforcement arrives with teeth, bulk sampling of frontier APIs gets metered, and proprietary weights keep some scarcity value. Or frontier capability compounds faster than a distiller can copy it, and the premium survives with no rule at all. Or the premium compresses toward open weight pricing and the frontier margin comes to rest on distribution and switching costs. Tan called the balance he wants "a tightrope" and said it "could result in the best possible outcome" [12].
My read is the third one. What would prove it wrong is an open weight model matching frontier benchmarks while frontier list prices hold, because that would locate the premium in distribution and reliability, which no amount of sampling can copy.
What to watch
- Whether the published text of the NSA, CISA and FBI advisory names companies or recommends controls on API access.
- Frontier API list prices at the next open weight release that matches them on published benchmarks.
- The outcome of The New York Times case against OpenAI and Microsoft, which sets what a lab owns of the corpus it trained on.