Product1 distinct publisher3 min readUpdated
A vendor panel at Supermicro's Open Storage Summit argues GPUs no longer set token cost alone. The useful corollary: chat, batch and agentic work each need their own capacity plan.
The Product Desk · Product desk

Compiled by The Product DeskSomething wrong?How this is made
A panel of infrastructure vendors put a claim on record this week that most capacity planning spreadsheets have not caught up with: GPU performance is still necessary, but storage latency, network bandwidth, data movement and power consumption increasingly determine the cost and speed of producing tokens [1]. If that holds, the unit economics of an inference fleet are no longer set by the accelerator invoice, and the three main production workload shapes each pull on a different constraint [2].
The workload split is the operationally useful part. Interactive chat prioritizes latency, batch inference emphasizes throughput, and agentic systems create expanding contexts [2]. Those are not three sizes of the same cluster; they are three different bottlenecks, which means one sizing model built on FLOPS will misprice at least two of them [15]. Ka Wai Leung of IBM's AI solutions product management put the sequencing plainly: "You need to understand what type of workload," he said. "Based upon the workload, you understand the characteristics of the workload, and you build your system behind it. That's how you scale" [3][4].
The access pattern shift is concrete. Training leans on high sequential throughput to feed models and finish checkpointing without idling GPUs, while inference puts the emphasis on low-latency random reads and writes, particularly where applications keep retrieving proprietary or recently updated data to supply context, according to Kioxia's Anders Graham [6]. Retrieval-augmented generation and multi-tenant AI factories raise the demand across the whole stack, according to Leung [3]. He also flagged the part that no drive fixes: enterprise data is scattered, and data gravity and sovereignty constrain how much of it can be ingested into an AI factory at all [8][9].
Then power. Data centers are running into limits on available energy and floor space, which turns efficiency into a capacity question rather than a cost line [10]. Kioxia's own measurement, comparing its BiCS8-based CM9 drives against the previous CM7 generation, showed a 76% improvement in random-read IOPS per unit of power and better than 100% on random write, according to Graham [11]. Taken at face value, that second figure means more than twice the random-write throughput per watt from a generational swap [12]. It is a vendor's number about its own product, on a vendor-run broadcast, so treat it as a hypothesis to test on your own traces [5].
The reference design being sold around this is an Nvidia HGX B300 compute environment with IBM Storage Scale Erasure Code Edition and Kioxia drives, using a fast tier for latency-sensitive work and an optional capacity tier for colder data, per Supermicro's William Li [13]. Li's pitch was consolidation of purchasing: rack integration, liquid cooling and storage from one vendor with partners attached [14]. That is a procurement argument, not an architecture argument, and it should be priced as one.
The item worth tracking is the KV cache work. Nvidia, IBM and Supermicro tested IBM Storage Scale as a shared KV cache, so previously computed context can be reused without holding all of it in GPU memory or system RAM [7]. That is the direct lever on agentic cost, where context keeps growing [2]. The published account of the test result breaks off before giving a figure [16], so the number to ask for is the cache-hit latency against local memory, and what share of tokens actually hit.
Follow any of these and your For You feed starts watching them — no settings page required.
Ranked by verification strength, evidence, and original report placement.
AI inference infrastructure is becoming a system-level challenge as organizations move generative and agentic applications into production. GPU performance remains essential, but storage latency, network bandwidth, data movement and power consumption increasingly determine the cost and speed of producing tokens.
Requirements vary by workload: interactive chat prioritizes latency, batch inference emphasizes throughput, and agentic systems create expanding contexts.
Retrieval-augmented generation and multi-tenant AI factories intensify demands across the stack, according to Ka Wai Leung, AI solutions product management at IBM Corp.
Leung said: "You need to understand what type of workload. Based upon the workload, you understand the characteristics of the workload, and you build your system behind it. That's how you scale."
Nvidia, IBM and Supermicro tested IBM Storage Scale as a shared KV cache, allowing previously computed context to be reused without keeping all cached data in limited GPU memory or system RAM.
Comparing its BiCS8-based CM9 drives with the preceding CM7 generation, Kioxia measured a 76% improvement in random read and over 100% improvement in random write, in input/output operations per second per unit of power, according to Graham.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Single sponsored vendor panel; all numbers self-reported
Every claim traces to one article covering one theCUBE panel of three suppliers, published under a paid media partnership. The quantitative items — KV cache throughput multiples, time-to-first-token, CM9 versus CM7 IOPS per watt — are asserted by the vendors that sell the parts, with no test configuration, workload definition or independent replication supplied. The qualitative framing (workload-specific bottlenecks, power limits) is plausible and internally consistent but is opinion from interested parties, and the ledger and article body even disagree about whether the cache figures were reported, which lowers confidence in the record itself.
Lab tests and a reference design; no disclosed customers
Observable adoption stops at supplier-side artifacts: a three-vendor lab test of Storage Scale as a shared KV cache, a Kioxia drive-generation measurement, and a marketed Supermicro reference architecture. No production deployment, named customer, deployment scale, shipment volume or usage disclosure appears, so real-world uptake of the described inference storage pattern is unevidenced.
Framing and multiples run ahead of the disclosed evidence
The headline thesis that the inference race has moved beyond GPUs, plus figures like '22 times more efficient' and doubled random-write IOPS per watt, are stated with more force than one sponsored panel of self-reported measurements can carry, and no customer outcome backs them. The gap is moderate rather than severe because the structural argument — per-workload bottlenecks, power and space ceilings, cache reuse to relieve GPU memory — is coherent, is partly qualified in the text (throughput fell under injected network noise), and the sponsorship is disclosed.
Paid media partnership plus three suppliers pitching their own stack
TheCUBE discloses it is a paid media partner for the Supermicro Open Storage Summit series, and each quoted speaker sells a component of the recommended architecture: IBM (Storage Scale), Supermicro (reference design, rack integration, liquid cooling) and Kioxia (CM9 SSDs), with Nvidia's HGX B300 and Spectrum-X as the named compute and network. The favorable numbers are produced and released by the same parties that monetize the conclusion, though the disclosure states sponsors lack editorial control.
Direction credible, specifics weakly grounded
Confidence is limited by single-source, sponsored provenance and by the absence of any independent measurement or deployment evidence; it is not lower because the article is explicit about attribution, quotes participants directly, discloses the commercial relationship, and reports at least one degradation result rather than only best-case numbers. The internal disagreement between the ledger note and the stored body about the cache figures is a further reason to hold the specific multiples loosely.
product
Marvell's $12.2bn warrant pays Google in Marvell stock, one $500m order at a time2 distinct publishers
invest
The AI moat is now a balance sheet, so price the financing and not the model1 distinct publisher
product
183 towns said no, so AI compute is heading offshore, underground, and into orbit1 distinct publisher
product
The shortage is the schedule: AI capacity relief is not a 2027 line item1 distinct publisher
Distinct publishers with included, body-backed reporting in this cluster.
1 article · August 19, 2026