Build1 distinct publisher3 min readUpdated
A presentation transcript published by InfoQ argues you cannot reason about inference economics without the chips, the power and the measured behaviour of the provider you rented from.
The Engineer · Build desk

Compiled by The EngineerSomething wrong?How this is made
Jordan Nanos, on the technical staff at the research firm SemiAnalysis, used a talk published in transcript by InfoQ to aim an argument at software teams: he says it is hard for him to use AI and feel he understands what is going on without understanding the chips, the data centers, the systems and the entire supply chain feeding them [1][2][4]. That matters because the number most engineering teams plan against is a published price per token, and Nanos's four-part framing puts almost every input to that number somewhere upstream of the API [11].
He framed the talk around one question: what is the choke point [10]. His answer starts with capital. On the chart he presented, the four large buyers - Google, Amazon, Meta and Microsoft - have revised 2026 spending up above analyst consensus, mostly on chips plus the data center capex around them, against a widely circulated figure of a trillion dollars going forward [15][16]. The constraint he pointed to is the other side of that spend: TSMC's quarterly wafer output, he said, is not growing exponentially the way the spending charts are, and the takeaway he drew is that TSMC is not keeping up with demand growth [18].
Treat that as directional rather than audited. Nanos noted the accelerator-spend chart he showed carries no y-axis, and the published transcript stops mid-sentence in the wafer discussion, so the volumes behind the conclusion are not in the material [17][20].
The part with the most immediate operational value is the measurement work. Nanos said ClusterMAX tests more than 100 cloud providers hands-on, hyperscalers and neoclouds alike, and publishes the findings for free [12]. A sister project, InferenceX, runs popular open-source models across NVIDIA and AMD hardware, with TPU, Trainium and several startup accelerators planned, and is also published free [13]. The reasonable inference from building a hands-on ranking of a hundred vendors is that those vendors are not interchangeable at the same nominal GPU and the same nominal price - though this transcript does not give the spreads, so anyone quoting a gap between providers is going beyond what is on the page here [12][13][20].
Fourth comes demand, which SemiAnalysis calls tokenomics: ChatGPT growth, Claude Code against Codex, and where in the stack the value actually lands [14]. That is the layer most software teams already argue about, and it is the layer furthest from the fab.
Worth registering the vendor's own position. SemiAnalysis sells research subscriptions covering data centers, chips, tokenomics and wafer fab equipment to investors and industry buyers, delivered with financial models and charts, and says its newsletter reaches 280,000 subscribers across three tiers [9][6]. Nanos joined last summer as employee 31; the firm now has more than 85 people, close to a tripling of headcount in roughly a year, and attends more than 100 conferences annually [7][6][8][19]. The free ClusterMAX and InferenceX data is useful and it is also a funnel [12][13][9].
What to watch: whether the 2026 capex actually lands above consensus for all four hyperscalers [15]; whether TSMC's quarterly wafer numbers, when published with axes, support the choke-point claim [18]; and whether InferenceX's TPU and Trainium results arrive in time to be compared against the NVIDIA and AMD baselines already there [13].
Follow any of these and your For You feed starts watching them — no settings page required.
Ranked by verification strength, evidence, and original report placement.
SemiAnalysis is described as a semiconductor and AI research firm and claims to be the number one Substack in the technology sector.
SemiAnalysis says its three-tier newsletter goes to 280,000 subscribers and that it has over 85 team members.
Nanos says he joined SemiAnalysis last summer as employee number 31.
SemiAnalysis says it generally attends 100-plus conferences a year.
SemiAnalysis's core business is selling research subscriptions to specific feeds covering data centers, chips, tokenomics and wafer fab equipment, to both investors and industry companies, delivered as emails with financial models and charts.
On the ClusterMAX project, SemiAnalysis tests hands-on all of the cloud providers in the industry, hyperscalers and neoclouds, over 100 of them, writes up its findings and gives out that research for free.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Single self-reported transcript, charts without axes
Everything rests on one InfoQ transcript of one speaker describing his own firm's paid research. The speaker explicitly notes the accelerator spend chart has no y-axis, wafer volumes are conveyed as chart colours by process node, and the text breaks off mid-sentence inside the TSMC discussion. Nothing here is corroborated by TSMC, NVIDIA, AMD, hyperscaler filings or a second publisher, so directional assertions are traceable but magnitudes are not checkable.
Programmes live and free, uptake self-reported only
There is real disclosed activity: ClusterMAX has hands-on tested over 100 clouds, InferenceX data is stated to be publicly available now at inferencex.com, and the research business claims 280,000 newsletter subscribers with 85-plus staff. But every figure is the vendor's own, no provider score, benchmark result, customer or downstream user of these rankings is named, and the hyperscaler CapEx and 3nm allocation observations carry no numbers, so measured uptake stays modest.
Confident choke-point verdict, axes withheld
The talk poses the bubble question and delivers a definite answer — TSMC wafer supply is 'effectively the key constraint' — while the supporting exhibits are deliberately unquantified and the transcript never reaches the tokenomics section that would connect fab capacity to token prices. The direction is plausible and internally consistent, and the free ClusterMAX/InferenceX work is understated rather than oversold, so the gap is moderate overstatement of certainty rather than fabrication.
Research vendor selling the models behind the thesis
The speaker's employer sells subscription feeds on exactly the topics of the talk — data centers, chips, tokenomics, wafer fab equipment — to investors and industry companies, and the paywalled institutional chart is shown with its axis removed. Free ClusterMAX and InferenceX output plus 100-plus conference appearances a year function as a funnel for that paid product, and the transcript contains no disclosure discussion. This describes commercial incentive structure, not inaccuracy.
One publisher, one speaker, truncated text
Low confidence in the magnitudes and moderate confidence only in what was said. There is a single publisher and a single primary voice, the strongest claims are unquantified by design, and the transcript is cut off before the sections that would test the thesis. Confidence would rise materially with TSMC or hyperscaler disclosure, published ClusterMAX scores, or a second outlet reporting the same wafer constraint independently.
product
Washington's secret AI test is coming for open weights, and release dates go with it2 distinct publishers
product
AMD borrows $4.75bn while sitting on $13bn, and the number matches its Anthropic promise1 distinct publisher
invest
The AI moat is now a balance sheet, so price the financing and not the model1 distinct publisher
build
1.5% of Hugging Face repos take 99.2% of downloads, and the ceiling is Chinese1 distinct publisher
Distinct publishers with included, body-backed reporting in this cluster.
1 article · August 18, 2026