Skip to content

Build1 publisher3 min readPublished

AMD's Venice white paper sets a GCC 16.1 build against Nvidia's GCC 15.2 Vera scores

The 20 percent per-core margin over Nvidia's Vera is the same claim AMD made in July. The detail now published behind it rests on a 256-core part down-cored to 96, a 600W budget, and two different major GCC releases.

The Engineer · Build desk

Photograph accompanying AMD's Venice white paper sets a GCC 16.1 build against Nvidia's GCC 15.2 Vera scores
Photo: tomshardware.com

What happened

  • AMD's new Venice white paper keeps its July claim that a 96-core high-frequency chip is around 20 percent faster than Nvidia's 88-core Vera in SPEC CPU 2026's Integer Rate test.
  • AMD's earlier SPEC runs gave the 96-core configuration the same 600W budget as the 256-core part, while the 96-core high-frequency SKU AMD sells tops out at 500W.
  • On the older July throughput slide, AMD puts the 256-core 9996 at 2.37 times the Intel Xeon 6980P and 2.24 times Vera.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • constraint The published comparison mixes GCC 16.1's contribution with the silicon's, and a shop standardised on GCC 15.2 has no supported path to the same margin.
  • cost Whoever budgets rack power pays the difference between the tested envelope and the rated cap of the part on the price list, and the margin was measured at the higher one.
  • decision Sizing a refresh off the agentic AI panel now requires asking AMD for its derived TPC-H and TPC-C harness, because those bars cannot be checked against published TPC results.
  • contradiction Tom's Hardware calls the mixed-compiler comparison bad practice while also reporting that GCC 16 pushes the compilation subtests down and other binaries up, so the bias runs in both directions across the table.

How much of the table transfers to your fleet depends on the compiler. AMD compiled its 96-core Venice subtest run with GCC 16.1 and set the results beside Vera scores Nvidia published from GCC 15.2 [5]. Michael Larabel of Phoronix has written up the difference between the two releases. Tom's Hardware's summary of it: GCC 16 takes longer to compile because it optimizes more, which is why the GCC and LLVM compilation subtests come in lower, and the binaries it emits then run faster by an amount that depends on flags, software and other factors [6]. On the method, Tom's Hardware wrote that "it's not best practice to compare benchmarks using two different compiler versions" [7].

The per-core margin transfers only if both sides are built with the compiler and flags you plan to ship and the 96-core part runs at the power the benchmark had. Isolating the silicon would take someone rebuilding Vera on GCC 16.1. On power, the tested budget is 20 percent above the cap of the high-frequency SKU AMD sells [10]. The per-core claim is also 20 percent. I do not think one number explains the other, but I would want the 500W run before signing anything.

The Stream comparison borrows its Vera side from Phoronix's initial controlled testing of the chip [11]. AMD's own side is a 9996 down-cored to 96 cores with a 600W budget, and it comes out about 18 percent ahead in total bandwidth and about 8 percent ahead per core [12]. Those two figures agree with each other: 96 cores against 88 is 9 percent more cores, and 1.18 divided by 1.09 is 1.08 [13].

The throughput slide is the oldest material in the paper, gathered in July with GCC 15.2 [14]. It also prices Nvidia against Intel, since 2.37 divided by 2.24 puts Vera about 6 percent ahead of the Xeon 6980P [23]. Standard practice on the intrate test is one copy per thread [16]. The white paper does not say what AMD ran, in the footnotes or anywhere else [16]. The generational number is AMD measured against AMD: AMD puts the 9996 about 78 percent faster than its own 192-core EPYC 9965 [17].

AMD ran the cloud workloads itself, covering database, Java and cryptography, with the Graviton5 results taken from an AWS cloud instance [18]. The agentic AI panel uses real benchmarks despite the labels on the chart: NGINX, the TPCx-AI kit, FAISS, TPC-H and TPC-C, and a replay of a multi-persona agent [19]. AMD says it derived the TPC-H and TPC-C workloads, so those bars cannot be reconciled with published TPC results [20]. In HPC, AMD extends its lead over Intel's Granite Rapids-AP flagship [22], and Intel's next-generation data center parts, Diamond Rapids, are due next year [21].

What to watch

  • Either vendor republishing Venice and Vera intrate scores built with the same GCC release, which would show how much of the margin is silicon.
  • A SPEC result for the 96-core high-frequency Venice SKU at its rated 500W cap.
  • Diamond Rapids shipping next year and resetting the Xeon 6980P baseline these charts are drawn against.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories