Product1 publisher3 min readPublished
Inference cost is now a storage and power problem, and it prices differently per workload
A vendor panel at Supermicro's Open Storage Summit argues GPUs no longer set token cost alone. The useful corollary: chat, batch and agentic work each need their own capacity plan.
The Product Desk · Product desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened
- AI inference infrastructure is becoming a system-level challenge as organizations move generative and agentic applications into production. GPU performance remains essential, but storage latency, network bandwidth, data movement and power consumption increasingly determine the cost and speed of producing tokens.
- Requirements vary by workload: interactive chat prioritizes latency, batch inference emphasizes throughput, and agentic systems create expanding contexts.
- Retrieval-augmented generation and multi-tenant AI factories intensify demands across the stack, according to Ka Wai Leung, AI solutions product management at IBM Corp.
- Leung said: "You need to understand what type of workload. Based upon the workload, you understand the characteristics of the workload, and you build your system behind it. That's how you scale."
- Leung, William Li (general manager of solution management at Super Micro Computer Inc.) and Anders Graham (senior director of SSD marketing and business development at Kioxia Holdings Corp.) spoke with theCUBE's Rob Strechay for the Supermicro Open Storage Summit interview series, broadcast on theCUBE, SiliconANGLE Media's livestreaming studio, with a disclosure noted.
Compiled by The Product DeskSomething wrong?How this is made
Why it matters
A panel of infrastructure vendors put a claim on record this week that most capacity planning spreadsheets have not caught up with: GPU performance is still necessary, but storage latency, network bandwidth, data movement and power consumption increasingly determine the cost and speed of producing tokens [1]. If that holds, the unit economics of an inference fleet are no longer set by the accelerator invoice, and the three main production workload shapes each pull on a different constraint [2].
The workload split is the operationally useful part. Interactive chat prioritizes latency, batch inference emphasizes throughput, and agentic systems create expanding contexts [2]. Those are not three sizes of the same cluster; they are three different bottlenecks, which means one sizing model built on FLOPS will misprice at least two of them [15]. Ka Wai Leung of IBM's AI solutions product management put the sequencing plainly: "You need to understand what type of workload," he said. "Based upon the workload, you understand the characteristics of the workload, and you build your system behind it. That's how you scale" [3][4].
The access pattern shift is concrete. Training leans on high sequential throughput to feed models and finish checkpointing without idling GPUs, while inference puts the emphasis on low-latency random reads and writes, particularly where applications keep retrieving proprietary or recently updated data to supply context, according to Kioxia's Anders Graham [6]. Retrieval-augmented generation and multi-tenant AI factories raise the demand across the whole stack, according to Leung [3]. He also flagged the part that no drive fixes: enterprise data is scattered, and data gravity and sovereignty constrain how much of it can be ingested into an AI factory at all [8][9].
Then power. Data centers are running into limits on available energy and floor space, which turns efficiency into a capacity question rather than a cost line [10]. Kioxia's own measurement, comparing its BiCS8-based CM9 drives against the previous CM7 generation, showed a 76% improvement in random-read IOPS per unit of power and better than 100% on random write, according to Graham [11]. Taken at face value, that second figure means more than twice the random-write throughput per watt from a generational swap [12]. It is a vendor's number about its own product, on a vendor-run broadcast, so treat it as a hypothesis to test on your own traces [5].
The reference design being sold around this is an Nvidia HGX B300 compute environment with IBM Storage Scale Erasure Code Edition and Kioxia drives, using a fast tier for latency-sensitive work and an optional capacity tier for colder data, per Supermicro's William Li [13]. Li's pitch was consolidation of purchasing: rack integration, liquid cooling and storage from one vendor with partners attached [14]. That is a procurement argument, not an architecture argument, and it should be priced as one.
The item worth tracking is the KV cache work. Nvidia, IBM and Supermicro tested IBM Storage Scale as a shared KV cache, so previously computed context can be reused without holding all of it in GPU memory or system RAM [7]. That is the direct lever on agentic cost, where context keeps growing [2]. The published account of the test result breaks off before giving a figure [16], so the number to ask for is the cache-hit latency against local memory, and what share of tokens actually hit.