Build1 publisherNot yet confirmed elsewhere3 min readPublished
AMD benchmarks Gorgon Halo against Intel in the week Nvidia's RTX Spark is expected
AMD published Gorgon Halo AI benchmarks days before Nvidia's expected RTX Spark launch, claiming a 1.1x to 32.2x lead over Intel's Core Ultra X9 388H. Until RTX Spark is measured, buyers weighing Gorgon Halo machines that cost upwards of $7,099 can compare the two only on memory.
The Engineer · Build desk

What happened
- Gorgon Halo supports up to 192GB of unified memory, while RTX Spark devices top out at 128GB, the same as the nearly identical GB10 chip in Nvidia's DGX Spark.
- AMD quotes a peak of 20 tokens per second on the 320-billion-parameter GLM 5.3 Flash using Unsloth's UD-IQ4_XS mixed quantization.
- In a press prebriefing AMD claimed tens of millions of AI PCs shipped, then clarified that it had shipped over half a million agentic PCs.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- constraint The AMD test box had three times the Intel system's memory, so its ratios blend chip speed with what fits in RAM and cannot be carried over to a 128GB RTX Spark.
- decision Teams whose models fit in 128GB cannot settle Gorgon Halo versus RTX Spark from the spec sheet; that purchase now waits on independent throughput tests.
- cost Long-context work pays for Gorgon Halo's larger models in speed, because AMD published peak rates and throughput falls as context length grows.
The ratios come from a rig with unequal memory. AMD ran its top-spec 192GB Ryzen AI Max+ Pro 495 against a Core Ultra X9 388H system with 64GB, on an Intel platform that supports up to 128GB [9]. The AMD machine had three times the memory [22]. AMD averaged multiple ComfyUI runs across various models and compared total throughput [8]. On that setup, a ratio folds chip speed and model fit into one number. Tom's Hardware called the 32.2x top end a clear outlier [26]. It found no model named Yuve on Hugging Face and said the gap could be an optimization issue, or a model too big to run on the Panther Lake machine [27].
AMD would argue it compared top-of-stack parts; Tom's Hardware says the two chips are not in the same class [7]. Gorgon Halo mainly competes with RTX Spark, and to a lesser extent with Apple's larger M-series SoCs [11]. Nvidia's part has no published performance results beyond a few questionable Geekbench leaks [6]. Its launch is expected at Microsoft's event on Wednesday, October 7 [3].
The token figures need the same care. GLM 5.3 Flash has 320 billion parameters but activates 18 billion per token [17]. AMD's peak on it is 20 tokens per second in Unsloth's UD-IQ4_XS format [15]. Tom's Hardware's own DGX Spark peak, 64 tokens per second, came from GPT-OSS 120B at 4-bit [16]. That model activates about 5 billion parameters per token [17]. GLM activates 3.6 times as many [24], so the 20 and the 64 measure different jobs.
Qwen 3.8 Flash Next is the closer match, with 5 billion active parameters per token. AMD quotes up to 42 tokens per second on it, using Unsloth's dynamic 4-bit quantization and multi-token prediction [1]. The last-gen Ryzen AI Max+ 395 ran GPT-OSS 120B at 56 tokens per second in Tom's Hardware's testing [16], 14 more [21]. Model, quantization and lab all differ, so the gap supports no conclusion about the new chip. For any of these numbers to transfer to a buyer's workload, the model, the quantization and the context length would all have to match.
AMD published peaks [15][1]. Tom's Hardware wrote that in both cases the "up to" "carries a lot on its shoulders," and that performance drops at higher context lengths [18].
Gorgon Halo is largely a refresh of Strix Halo [10]. It sells into a category Tom's Hardware describes as far smaller than some of the RTX Spark hype suggests [19]. AMD told a press prebriefing it had shipped "10s of millions" of AI PCs, then clarified the figure as "over half a million" agentic PCs [12]. The number shrank by a factor of at least 20 between sentences [25].
Memory is the one axis with public numbers on both sides, and the matchup so far has turned mainly on it [20]. Gorgon Halo supports 192GB against RTX Spark's 128GB ceiling [13], 64GB more [23]. That capacity runs larger models locally, at lower performance [14]. I think the memory spec is the only part of AMD's pitch a buyer can check against Nvidia this week. A team whose models fit in 128GB is choosing on throughput, and that comparison arrives only with independent RTX Spark results. For that team I would hold an order for a machine that can cost upwards of $7,099 [5] until those results are out. A team that needs more than 128GB already has its answer in the spec sheet [13].
What to watch
- Independent RTX Spark throughput results at matched model, quantization and context length after the expected launch at Microsoft's October 7 event.
- Whether AMD publishes Gorgon Halo token rates across context lengths, beyond its peak 'up to' figures.
- A rerun of AMD's ComfyUI comparison with the Intel system at its 128GB maximum, to separate chip speed from model fit.