Product1 distinct publisher3 min readPublished
A CNCF walkthrough puts the money on accelerator utilisation rather than serving throughput, and argues Kubernetes now supplies most of the parts, with the isolation half of the job still sitting on the platform team's desk.
The Product Desk · Product desk

Compiled by The Product DeskSomething wrong?How this is made
Start with the arithmetic, because it carries the case. A pod that requests `nvidia.com/gpu: 1` and then uses a tenth of the card is sitting on the other nine tenths [1]. Ten workloads of that shape could in principle share one accelerator; the whole-device request model gives them one each [2]. A faster model server does not touch this, and the CNCF post says so directly: what is needed is a stack that allocates accelerators without stranding or endangering capacity, and isolates tenants well enough that packing them together holds [11].
Here is what platform teams tell themselves: the accelerator problem is a scheduling problem, so a version bump fixes it. Here is what the drawing actually shows. Most layers are Kubernetes native or CNCF projects, with a few OSS and vendor pieces at the points that decide density and separation, including NVIDIA's MIG, vCluster and Dynamo [12]. Two years ago these same teams already had containers, RBAC, autoscaling and policy working, and no clean answer for accelerators or for keeping tenants apart on one node [16]. The parts arrived separately, from different owners. NVIDIA's framing, "infrastructure for the full AI lifecycle, from data preparation through training, fine-tuning, and high-volume inference", describes the scope rather than the plumbing [2].
The provisioning layer is where the honesty lives. Before a node counts, something inventories its GPUs, checks ECC state, records InfiniBand GUIDs and NIC MACs, network boots an image with the driver, CUDA and NCCL baked in, then applies a BIOS profile chosen for baseline, performance or confidential compute [13]. Burn-in runs it under load, an NCCL test proves the GPUs talk at full bandwidth, and the result lands in a source of truth such as NetBox, which also holds the IP assignments [14]. Retirement runs the same loop backwards: disks wiped, the remote-management login reset, the node handed back clean [15]. A fleet count and a validated fleet count are two different numbers, and only one of them can be scheduled.
The forcing function worth putting on the GPU intake form has two axes: where the tenant's isolation requirement comes from, and how much of a device the workload fills. A contractual or regulatory boundary plus a full-device job gets its dedicated block, with the idle hours priced back to the tenant that required them, since that default is safe under strict trust and wastes most of the hardware [9]. A contractual boundary with a small job goes to a separate pool running the confidential-compute BIOS profile, because a slice does not soften a boundary. A habitual boundary with a small job is the one cell where MIG partitions under a virtual control plane earn their keep. A habitual boundary with a full-device job belongs in the shared pool, with quotas doing work they already do well.
Naming the origin of the boundary is the uncomfortable part, because "we have always had our own cluster" arrives on a capacity request looking exactly like a compliance requirement. What settles the argument a quarter later is validated device-hours per tenant, split into allocated and used, because that split prices your tenancy policy instead of your hardware.
Ranked by verification strength, evidence, and original report placement.
Accelerators are the dominant capital expense in the building, and the metric that decides whether that spend pays off is utilisation, not a peak tokens-per-second number from a single run.
The CNCF post describes an AI factory as a pool of GPUs many teams draw from at once: one team fine-tuning, another serving inference, a third running evaluations, all on the same accelerators.
NVIDIA frames an AI factory as "infrastructure for the full AI lifecycle, from data preparation through training, fine-tuning, and high-volume inference".
In an enterprise this means one fleet, many teams, and different quotas, policies and trust boundaries layered on top.
SemiAnalysis's ClusterMAX scores GPU cloud providers on security, networking, storage, reliability and support rather than raw throughput.
ClusterMAX's security criteria reward hard per-tenant isolation, down to per-tenant Kubernetes clusters and DPU-based isolation, while flagging weak boundaries such as putting many tenants on one cluster.
Follow any of these and your For You feed starts watching them — no settings page required.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Single-source architectural argument with a few checkable anchors
All material comes from one CNCF blog post. Its checkable anchors are real - DRA's GA status in Kubernetes 1.34, HAMi's CNCF Incubating status, and ClusterMAX's published scoring dimensions - and its provisioning walkthrough is specific enough to audit. But the load-bearing economic claims (utilisation, not throughput, decides payoff; dedicated GPU blocks waste most of the fleet) carry no measurements, and the field pattern of operators refusing demand is unattributed, so evidence sits below the midpoint.
Upstream primitives shipped, real-world fleet uptake undisclosed
There is genuine upstream adoption signal: DRA has reached GA in a shipped Kubernetes release and HAMi has CNCF Incubating status, and the post enumerates a concrete toolchain per layer. What is absent is any deployment-side evidence - no named operators running the assembled stack, no cluster counts, no utilisation gains, and the vendor provisioning layer is described generically. Adoption is therefore real at the project level and unquantified at the fleet level.
Mildly overstated: strong economic framing, no measured outcomes
The framing is more confident than the evidence supports - utilisation is declared the deciding metric and an allocation-plus-isolation stack is declared the fix, with no measured utilisation improvement anywhere in the piece, and the publisher's own project ecosystem is the recommended answer. The gap is small rather than large because the post repeatedly self-limits: DRA does not fractionalise GPUs, MIG's use as a boundary between hostile tenants is called contested, and whole-GPU-per-tenant is presented as the conservative default.
Foundation blog advocating its own ecosystem, plus an undisclosed commercial mention
The publisher is CNCF and the piece's conclusion is that most layers of the answer are Kubernetes-native or CNCF projects, including CNCF Incubating HAMi - a clear alignment between argument and institutional interest. The text also presents a build-versus-assemble choice and names a commercial provisioning option (vMetal) alongside Metal3/Ironic without any affiliation statement in the supplied body. This is normal ecosystem advocacy rather than concealed promotion, so the score is elevated but not extreme.
Moderate-low: one publisher, verifiable technical spine, unverified economics
Confidence is limited by having a single publisher with no corroboration or challenge in the cluster, and by the supplied body being truncated mid-section. It is supported by the fact that the technical spine - DRA GA status, the device-plugin whole-GPU behaviour, provisioning and validation steps, ClusterMAX criteria - is stated precisely enough to check independently, while the economic and field claims are not.
build
Shadow engines cut LLM restart from 283 seconds to 7.3, and change what headroom is for1 distinct publisher
build
A benchmark that replays real agent sessions gives back less of the generational win2 distinct publishers
build
Nvidia's Groq-derived LPX rack posts 3,431 tokens/sec on 128GB of SRAM1 distinct publisher
build
SemiAnalysis to software teams: your token cost starts at the fab, not the price list1 distinct publisher
Distinct publishers with included, body-backed reporting in this cluster.
1 article · August 27, 2026