OpenAI has contracted 750MW of Cerebras wafer-scale inference capacity, delivered in three 250MW segments by the end of 2028. The dates are contractual, but neither company has published pricing, latency figures or a plan for which customers get access.
Reality
- Evidence45
- Adoption25
- Hype gap+10
- Incentives45
- Confidence40
Cerebras says splitting inference stages across chip types gave 5x more throughput from the same number of its systems without slowing token generation. Because the count covers only Cerebras hardware, the figure does not yet show what a mixed-chip fleet costs per unit of work.
Reality
- Evidence30
- Adoption20
- Hype gap+30
- Incentives75
- Confidence40
OpenAI's Ultrafast tier runs GPT-5.6 Sol up to 14 times faster than standard, at as many as 750 output tokens a second, for a limited preview group. With no price published, teams can prototype real-time features on it but cannot yet budget a production launch.
Reality
- Evidence35
- Adoption10
- Hype gap+15
- Incentives50
- Confidence35
Gimlet Labs plans 100 MW of Cerebras-powered inference capacity, pairing wafer-scale chips with GPUs for different phases of each request. For buyers, the quoted 3,000 tokens per second is a company target until measured results appear.
Reality
- Evidence40
- Adoption15
- Hype gap+40
- Incentives70
- Confidence40
Bloomberg reports the London startup raised $100 million in seed money to route each task to the cheapest workable model-chip pairing. The wager is that heterogeneity outlasts convenience.
Reality
- Evidence55
- Adoption12
- Hype gap+35
- Incentives60
- Confidence60
The accelerator is in full production and the headline number is a single-request generation rate at 100,000 tokens of context. That is a different purchase order than throughput.
Perspective Coverage
3 publishers
- Builder
- Builder 48%
- Operator
- Operator 25%
- Investor
- Investor 27%
Reality
- Evidence55
- Adoption20
- Hype gap+35
- Incentives80
- Confidence60
Jalapeno's lead is measured against last generation and the volumes are tiny, but a first-pass ASIC clearing Nvidia, AMD and Google parts reprices the design barrier, not the supply chain.
Perspective Coverage
7 publishers
- Builder
- Builder 30%
- Operator
- Operator 24%
- Investor
- Investor 46%
Reality
- Evidence50
- Adoption8
- Hype gap+40
- Incentives65
- Confidence55
DigitalOcean says step-by-step reasoning is billed as output and invisible by design. The 90% figure it cites comes from a paper that estimates hidden token counts, so moving it onto your own invoice takes a matching task mix.
Reality
- Evidence32
- Adoption26
- Hype gap+38
- Incentives88
- Confidence42
In IEEE Spectrum's account of the 2026 compute market, reasoning and agentic workloads have pushed serving toward memory-heavy silicon, and Amazon now runs a single inference job across two vendors' chips.
Reality
- Evidence47
- Adoption55
- Hype gap+24
- Incentives66
- Confidence46
Positron says the money covers a TSMC N3P tapeout by the end of 2026, a new 2MW engineering data center and a production ramp in the second half of 2027. The larger of the round's two tranches is described as up to $500 million.
Reality
- Evidence55
- Adoption35
- Hype gap+32
- Incentives72
- Confidence62
A dev.to harness put five coding tasks through four APIs and every run passed its verifier on the first attempt, so the only thing left to compare is clock time, where one trial per cell sets a fragile order.
Reality
- Evidence34
- Adoption22
- Hype gap+35
- Incentives40
- Confidence55
The gate on recursive self-improvement is how long one loop takes, and experiments, training runs and new chips all take real time. That makes capability curves fast but bounded, which changes what you sequence first.
Publishers:exponentialview.co
Reality
- Evidence34
- Adoption56
- Hype gap+18
- Incentives58
- Confidence42
More memory on a Cerebras wafer arrives with CS-6, two generations out. Everything shipping before then, Nexus included, is rack engineering around the SRAM budget the wafer already has, and long contexts are where that budget hurts.
Reality
- Evidence42
- Adoption33
- Hype gap+34
- Incentives72
- Confidence41
Forge counts seven of twenty VC-backed debuts trading above their IPO price. Figma, up 250% on day one, is 84.3% below its first-day close.
Publishers:forgeglobal.com
Reality
- Evidence44
- Adoption58
- Hype gap+32
- Incentives86
- Confidence47
Forge's data puts today's AI cohort at $100 billion in about five years, against 16 for the SpaceX generation. The listing increasingly looks like an exit for someone else.
Publishers:forgeglobal.com
Reality
- Evidence42
- Adoption55
- Hype gap+34
- Incentives82
- Confidence44
The first third-party benchmark of the LP30 rack came in at roughly four times the next-fastest public endpoint, measured one request at a time on a model small enough to fit.
Reality
- Evidence58
- Adoption20
- Hype gap+32
- Incentives74
- Confidence55
The Nexus rack, not the WSE-3, is now the thing Cerebras ships. That makes upgrade cadence the number buyers should price, and 10,000 tokens per second a target rather than a plan input.
Reality
- Evidence56
- Adoption32
- Hype gap+34
- Incentives76
- Confidence62
Morgan Stanley closed its purchase of the secondaries marketplace in January, and the private shares changing hands there increasingly sit behind lenders who get paid first.
Reality
- Evidence34
- Adoption44
- Hype gap+20
- Incentives78
- Confidence46
The Information reports a valuation more than 50% above the last round, with a licensing arrangement considered alongside it. Buyers of AI search should read the roadmap as supplier-driven.
Reality
- Evidence34
- Adoption48
- Hype gap+42
- Incentives82
- Confidence52
Hardware from the largest acquisition in Nvidia's history is in production and headed to a neocloud this year. The market sold it anyway, two days before earnings.
Reality
- Evidence58
- Adoption28
- Hype gap+24
- Incentives74
- Confidence61
Earlier coverage
- One Anthropic order, six times the price: Fractile's $6.5B mark arrives two years before its chips
Invest · August 20, 2026 · 3 publishers
- Callosum's $100m seed is a 10x on February, and a UK state fund's first cheque
Invest · August 20, 2026 · 3 publishers
- Cerebras's CS-4 is three old wafers in a new rack: price the packaging, not the silicon
Product · August 19, 2026 · 2 publishers
- WhiteFiber pays $60m for two Unifi plants, and buys the shell rather than the site
Product · August 17, 2026 · 1 publisher
- Anthropic's IPO case rests on a 2028 number, and your roadmap sits inside it
Product · August 17, 2026 · 1 publisher
- 750 tokens a second moves agents from overnight batch to interactive latency
Security · August 15, 2026 · 1 publisher
- The shortage is the schedule: AI capacity relief is not a 2027 line item
Product · August 15, 2026 · 1 publisher
- Developer habit, priced at $965B: what Anthropic's run actually proves
Build · August 15, 2026 · 1 publisher
- Kog's pitch: the cheapest inference upgrade is the H200s you already bought
Product · August 15, 2026 · 1 publisher
- Three neoclouds, one pattern: revenue up, losses up, capex far ahead of both
Product · August 14, 2026 · 2 publishers
- OpenAI puts latency on the price list: 750 tokens/sec, gated by workload fit
Build · August 14, 2026 · 3 publishers
- Speed becomes a SKU: OpenAI and Google put a separate price on latency
Invest · August 14, 2026 · 3 publishers