Build1 publisher3 min readPublished
SemiAnalysis to software teams: your token cost starts at the fab, not the price list
A presentation transcript published by InfoQ argues you cannot reason about inference economics without the chips, the power and the measured behaviour of the provider you rented from.
The Engineer · Build desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened
- Jordan Nanos says he is on the technical staff at SemiAnalysis and primarily works on a project called ClusterMAX.
- The material is a transcript of a presentation titled 'From Fab To Token - The State Of The Market', published by infoq.com.
- Nanos describes himself as a practitioner and a hardware guy who uses models and develops software in service of hardware analysis, and says he cares about testing new GPUs, measuring performance and testing cloud providers.
- Nanos: 'it's really hard for me to use AI and feel like I really understand what's going on without having any understanding of the chips, the data centers, the systems, even the entire supply chain that goes into it.'
- SemiAnalysis is described as a semiconductor and AI research firm and claims to be the number one Substack in the technology sector.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
Jordan Nanos, on the technical staff at the research firm SemiAnalysis, used a talk published in transcript by InfoQ to aim an argument at software teams: he says it is hard for him to use AI and feel he understands what is going on without understanding the chips, the data centers, the systems and the entire supply chain feeding them [1][2][4]. That matters because the number most engineering teams plan against is a published price per token, and Nanos's four-part framing puts almost every input to that number somewhere upstream of the API [11].
He framed the talk around one question: what is the choke point [10]. His answer starts with capital. On the chart he presented, the four large buyers - Google, Amazon, Meta and Microsoft - have revised 2026 spending up above analyst consensus, mostly on chips plus the data center capex around them, against a widely circulated figure of a trillion dollars going forward [15][16]. The constraint he pointed to is the other side of that spend: TSMC's quarterly wafer output, he said, is not growing exponentially the way the spending charts are, and the takeaway he drew is that TSMC is not keeping up with demand growth [18].
Treat that as directional rather than audited. Nanos noted the accelerator-spend chart he showed carries no y-axis, and the published transcript stops mid-sentence in the wafer discussion, so the volumes behind the conclusion are not in the material [17][20].
The part with the most immediate operational value is the measurement work. Nanos said ClusterMAX tests more than 100 cloud providers hands-on, hyperscalers and neoclouds alike, and publishes the findings for free [12]. A sister project, InferenceX, runs popular open-source models across NVIDIA and AMD hardware, with TPU, Trainium and several startup accelerators planned, and is also published free [13]. The reasonable inference from building a hands-on ranking of a hundred vendors is that those vendors are not interchangeable at the same nominal GPU and the same nominal price - though this transcript does not give the spreads, so anyone quoting a gap between providers is going beyond what is on the page here [12][13][20].
Fourth comes demand, which SemiAnalysis calls tokenomics: ChatGPT growth, Claude Code against Codex, and where in the stack the value actually lands [14]. That is the layer most software teams already argue about, and it is the layer furthest from the fab.
Worth registering the vendor's own position. SemiAnalysis sells research subscriptions covering data centers, chips, tokenomics and wafer fab equipment to investors and industry buyers, delivered with financial models and charts, and says its newsletter reaches 280,000 subscribers across three tiers [9][6]. Nanos joined last summer as employee 31; the firm now has more than 85 people, close to a tripling of headcount in roughly a year, and attends more than 100 conferences annually [7][6][8][19]. The free ClusterMAX and InferenceX data is useful and it is also a funnel [12][13][9].
What to watch: whether the 2026 capex actually lands above consensus for all four hyperscalers [15]; whether TSMC's quarterly wafer numbers, when published with axes, support the choke-point claim [18]; and whether InferenceX's TPU and Trainium results arrive in time to be compared against the NVIDIA and AMD baselines already there [13].