Build1 distinct publisher3 min readPublished
One SKU, one compute die, 1.2 TB/s of LPDDR5X, and benchmark wins over a 96-core EPYC. The comparison with AMD's Venice is being set on Nvidia's terms rather than on core counts.
The Engineer · Build desk

Compiled by The EngineerSomething wrong?How this is made
Bandwidth per core is the number the slides leave for the reader to work out. Divide 1.2 TB/s by 88 cores and each core sits behind roughly 13.6 GB/s of memory bandwidth [17]. That figure, not the core count, is what a monolithic compute die and an LPDDR5X subsystem are built to protect [3][4].
The workloads Nvidia chose to publish reward cores. Kernel compilation and parallel headless browser instances both scale out, and Nvidia is running them against a part with eight more cores than Vera has [7][8][16]. Normalise the claims per core and the gap widens: 22 percent faster on a native AArch64 kernel build at 88 cores against 96 is about 33 percent more throughput per core, and the 24 percent browser-scaling win is about 35 percent [18][19]. Those are per-core and per-byte numbers dressed as socket numbers.
The cross-compile result is the one operators should read twice. Vera's advantage drops from 22 percent on a native AArch64 target to 14 percent when the target is x86 [8], which means roughly 7 percent of the win goes back to the toolchain when the fleet still builds x86 artefacts [20]. Agentic build chains that emit binaries for x86 hosts pay that back on every job.
Spatial multithreading is aimed at variance rather than peak. Nvidia splits core resources across two pipelines while still allowing data and cache to move between threads [10], and the SPEC CPU 2017 intrate demonstration measures a core alone, then with a second thread live [11]. Traditional SMT time-slices the resources and leaves gaps at branch prediction and decode; Nvidia's version still has threads contending, but degrades in a predictable way [12]. For a serving fleet, predictable degradation is worth more than a higher single-thread ceiling, because tail latency is what you have to provision against.
Which makes the missing comparator awkward. Nvidia did not say which "traditional CPU" the noisy-neighbour chart is measured against [11], so the single result supporting the architecture's most distinctive feature cannot be checked by anyone else. The 4.5x figure attached to the browser workflow deserves the same discipline: it comes from stripping GUI rendering, fonts and media decoding out of the browsing path [9], which is a property of the workload rather than of the silicon, and any CPU running the same trimmed path inherits it.
On Venice, the disclosed numbers do not exist. Tom's Hardware points to earlier coverage of how Vera stacks up against AMD's next-generation parts, but the comparisons Nvidia brought to Hot Chips 2026 are against the shipping 96-core EPYC 9655P [16]. So the like-for-like argument is being fought on ground Nvidia picked: one SKU [3], one compute die instead of compute chiplets [4], memory bandwidth, and consistency under load. Nvidia's own framing concedes the weakness in that ground, since it acknowledges agentic AI has no settled benchmark and that tasks such as compilation are proxies for chains that are long and inconsistent to measure [5][6]. A buyer comparing sockets in 2026 will be comparing a fixed 88-core part against a family, using benchmarks that stand in for the workload rather than being it.
Ranked by verification strength, evidence, and original report placement.
Nvidia says Vera is designed specifically for agentic AI workloads, a category still being defined in terms of performance benchmarking.
Many CPU-intensive tasks such as code compilation serve as proxies for agentic workloads, but measuring performance across a full agentic chain is complex and inconsistent.
Spatial multithreading separates core resources across two pipelines, though data and cache can move between threads as needed.
With traditional SMT, resources are time-sliced between threads, leaving gaps between branch prediction and decode; Nvidia's spatial multithreading still has threads competing for in-core resources but handles neighbouring demand deterministically, giving a more consistent per-core downturn.
Nvidia continued its Vera CPU disclosures at Hot Chips 2026, adding detail on spatial multithreading, the memory subsystem, and the workloads it is targeting.
Vera is the first CPU with a custom Nvidia core, following Grace, which used a stock Arm design.
Follow any of these and your For You feed starts watching them — no settings page required.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Detailed but entirely vendor-supplied and single-sourced
The architectural disclosures are specific and internally consistent (single 88-core SKU, monolithic die, spatial multithreading mechanism, SOCAMM2 LPDDR5X at 1.2 TB/s), which lifts the floor. But every performance number originates from Nvidia's own slides at a conference talk, one publisher reports them, the noisy-neighbour comparator is unnamed, and the reporting itself notes the bandwidth-per-watt chart measures power against peak bandwidth. No independent measurement exists in the supplied material.
Pre-deployment disclosure, no third-party usage
Adoption evidence stops at Nvidia's own conference disclosure and vendor benchmarks: a stated single-SKU shipping configuration and platform memory design. The supplied source names no customer, cloud deployment, order, or independent evaluation, so the only observable activity is vendor communication ahead of availability.
Vendor framing runs ahead of verifiable results
Nvidia's slides call agentic AI the 'most complex computing workload in history' and a 4.5x browsing-workflow figure is presented for a workload class the reporting concedes has no consistent measurement. The supporting numbers are modest single-digit-to-low-double-digit socket-level wins over a current-generation EPYC, produced by the vendor, against a comparator chosen by the vendor, with the SMT comparison partner unnamed. That gap between rhetorical framing and demonstrated margin is positive but bounded, because the reporting flags the caveats rather than amplifying them.
Vendor-controlled disclosure with clear competitive stake
Nvidia authored the benchmarks, chose the comparator, defined the workload proxies and set the framing for a product entering AMD's data-center CPU territory; the bandwidth-per-watt presentation is flagged by the reporting as potentially favourable to Nvidia. Distribution incentives are also visible: the publisher notes the premium article was temporarily free to widen Hot Chips readership. Nothing here is hidden, but the incentive concentration on the source side is high.
Solid on architecture, weak on performance and uptake
Confidence is bounded by a single publisher relaying a single vendor. The architectural and platform facts are reported precisely and are unlikely to be wrong in substance; the performance claims and the agentic-workload rationale depend on unreproduced vendor data, and adoption is essentially unobserved. Derived per-core arithmetic is arithmetically sound but inherits the underlying claims' uncertainty.
build
Samsung puts MAC trees in every LPDDR5X bank because HBM costs too much1 distinct publisher
build
China's accelerator swap makes Cambricon supply, not export policy, your ship-date risk1 distinct publisher
product
Waymo opens the trunk: a 5nm ASIC, a quadrillion ops, and a supplier list rivals can price3 distinct publishers
build
IBM put an Arm decoder in a mainframe core, and only half the consolidation pitch has a native path1 distinct publisher
Distinct publishers with included, body-backed reporting in this cluster.