Build1 distinct publisher2 min readPublished
The bandwidth gap is 4.4x and the capacity gap is 4x, which is why these two boxes are not really competing. One decides whether a model fits; the other decides whether it is usable.
The Engineer · Build desk

Compiled by The EngineerSomething wrong?How this is made
Take the configuration both vendors point at and divide. A 200B-parameter model at 4-bit weights is roughly 100GB of memory [1], and decode produces one token at a time, with a full pass over the weights for each one [7]. That makes the memory bus a hard ceiling: 273GB/s across 100GB is about 2.7 tokens per second on DGX Spark [2], while 1.2TB/s across the same weights is about 12 [3]. Both figures sit before KV cache traffic and everything else a runtime does, so read them as the rate no benchmark will beat. NVIDIA's advertised support for up to 200 billion parameters [8] is a claim about what fits, and it holds. The quoted petaFLOP of FP4 [15] is compute for prefill and training, not for the token you are waiting on.
Price per gigabyte then crosses over. Spark is $4,699 for 128GB, or $36.71 per gigabyte [4]. The base Mac Studio is $5,499 for 96GB, or $57.28 [5], the worst ratio in the comparison. Apple's 256GB step costs $4,000 on top of base [4], so that machine lands at $9,499 and $37.10 per gigabyte [6], within about one percent of Spark [7] with the bandwidth advantage attached. The 512GB tier has no price and does not ship until late October [3], which is the part a purchase order cannot act on. Note also what the headline number actually is: 1.2TB/s is a 50 percent step over M3 Ultra's 800GB/s [9], so the ceiling moves 50 percent and the arithmetic above moves with it.
One caution on all of it. This comparison comes from a single dev.to analysis whose author states he has used neither machine and worked from Apple's newsroom, NVIDIA's datasheet and third-party DGX Spark reviews [13]. Spec sheets are precisely where sustained behaviour hides, which is what the Spark throttling reports were about.
So the jobs separate. If the work is generating tokens from a large model at a rate a person will sit through, the bus decides and there is one machine here. If the work is tuning, or anything that assumes CUDA, the spec sheet stops mattering: the inference servers people actually deploy were born on CUDA, and MLX remains Apple-only and slower to get day-one support for new open models [14]. Both boxes run off a standard wall outlet [16], so the constraint that used to settle this argument at home is gone, and the two memory numbers are what is left.
Ranked by verification strength, evidence, and original report placement.
Apple announced a new Mac Studio with M5 Ultra offering up to 512GB of unified memory at 1.2TB/s of memory bandwidth, starting at $5,499 with 96GB.
Mac Studio M5 Ultra starts at $5,499 with 96GB of unified memory.
In the decode phase of LLM inference the model generates one token at a time and every token requires a full pass over the weights; once the model is loaded, inference is almost entirely a memory-bandwidth problem.
NVIDIA's published numbers show a Llama 3.3 70B QLoRA fine-tune hitting a peak of 5,079 tokens per second of training throughput on DGX Spark.
Serious inference servers such as vLLM and TensorRT-LLM were born on CUDA, while Apple's MLX ecosystem is Apple-only, smaller, and slower to receive day-one support for new open models.
NVIDIA quotes up to 1 petaFLOP of FP4 for DGX Spark.
Follow any of these and your For You feed starts watching them — no settings page required.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Vendor specs plus third-party DGX benchmarks, no hands-on and no primary links
The cluster contains exactly one source, a developer blog post whose author discloses he has used neither machine. Hard numbers on the DGX Spark side are traceable to NVIDIA's datasheet and to named independent reviews (Hardware Corner, Level1Techs), which is genuine third-party grounding. The Apple side rests entirely on relayed newsroom specs plus bandwidth arithmetic, with no reviewer units, no measured decode figures, and an internal inconsistency in the M3 Ultra baseline (800GB/s vs 819GB/s). The reasoning about bandwidth-bound decode is sound and self-consistent, but nothing in the cluster independently verifies the M5 Ultra configuration or price.
One box shipping and benchmarked, the other announced with its key config unpriced
DGX Spark is described as shipping now at a known price with independently benchmarked performance and post-launch field reports, which is real deployment evidence. The Mac Studio M5 Ultra is at announcement stage: the 96GB base has a price, the 256GB step has an upgrade cost, the 512GB configuration that carries the whole capacity argument is unpriced and dated late October, and reviewers have no units. There are no user counts, unit volumes, or organizational deployments in the cluster for either machine.
Debunks vendor headlines, then leans on its own unmeasured ceilings
The article is itself corrective on vendor hype: it discounts NVIDIA's 'up to 1 petaFLOP' FP4 headline against measured generation rates of 38.6 tok/s and 5.4 tok/s, and flags that Apple's '4.3x faster AI performance' compares against the two-generation-old M3 Ultra and covers prompt processing rather than decode. That pushes the gap toward zero. It stays positive because the comparison's decisive numbers, a 4.4x bandwidth advantage translating 'almost directly' into tokens per second and a ~12 tok/s Mac ceiling, are unmeasured arithmetic on hardware nobody has reviewed, priced against a 512GB configuration Apple has not priced, and because the cited DGX measurements show real decode falling well short of naive bandwidth math.
Vendor marketing materials as primary inputs, offset by an explicit no-hands-on disclosure
Two of the three input streams are vendor-controlled (Apple newsroom, NVIDIA datasheet), and both vendors have a direct commercial interest in the desk-side local-AI framing the article adopts. Mitigating factors: the author discloses up front that he has used neither machine, names his sources, states his buying lens, and actively discounts both vendors' headline claims using third-party reviews. The cluster discloses no vendor relationship, sponsorship, or affiliate arrangement in either direction, so this score reflects only the sourcing incentives that are visible.
Single publisher, one-sided measurement, decisive SKU unpriced
Confidence is limited by structure rather than by internal quality. There is one source and one publisher, so no cross-publication corroboration is possible; measured performance exists for only one of the two machines; the Apple configuration that anchors the capacity argument is unpriced and unshipped; and the source contradicts itself on the M3 Ultra bandwidth baseline. The spec-level facts about DGX Spark and the qualitative software-ecosystem contrast are the parts of this story that can be held with reasonable confidence.
build
Your 2026 GPU Decision Is Arithmetic: Bytes Per Parameter, Times Parameters, Plus Cache1 distinct publisher
build
Inco AI's DFlash 2: 21% longer accepted drafts for 1.3% latency and 18.5M parameters1 distinct publisher
build
Base Compute hands kernel tuning to agents; the carryover claim is the unmeasured part1 distinct publisher
leadership
Apple's $18,299 Mac Studio versus a $200 subscription: 91 months to break even1 distinct publisher
Distinct publishers with included, body-backed reporting in this cluster.
dev.to
1 article · August 25, 2026