Skip to content

Build1 publisher2 min readPublished

The 1.2 TB/s Mac Studio asks you to pick your model size before you buy it

Tom's Hardware puts the M5 Ultra at up to 1.2 TB/s of memory bandwidth over as much as 256GB of soldered unified memory, so the capacity that decides which weights stay resident is settled at checkout.

The Engineer · Build desk

Photograph accompanying The 1.2 TB/s Mac Studio asks you to pick your model size before you buy it
Photo: tomshardware.com

What happened

  • Apple's M5 Ultra is its first quad-die chip, built by connecting two dual-die M5 Max SOCs over the company's UltraFusion link.
  • Apple rates the chip's unified memory bandwidth at up to 1.2 TB/s, which it says allows large language models to be stored and run entirely on the device.
  • The unit Tom's Hardware tested sells for $12,299 with the top-end M5 Ultra, 256GB of unified memory and 4TB of storage, while configurations start at $2,499 with an M5 Max.
  • The RAM is soldered to the board and Apple does not sell replacement SSDs, so nothing inside the machine can be changed after the sale.
  • Memory options currently reach 256GB, with a 512GB configuration due later.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • decision The memory tier is a one-time decision. A team whose resident weights outgrow 256GB buys a second machine, because the soldered RAM rules out adding to the one it has.
  • cost Reaching the 256GB, 4TB configuration costs $9,800 more than the entry machine, and that capital is committed before the first token is served.
  • contradiction The verdict rests its local-AI recommendation on a throughput spec, so a buyer weighing the box against rented GPU time has to measure their own model to learn what the bandwidth yields.
  • constraint With few PCIe connections and fixed storage, Thunderbolt 5 and the 10Gb Ethernet jack are the only routes to more capacity once the box is on the desk.

Divide the bandwidth by the capacity. 1.2 TB/s over 256GB is about 4.7 full reads of memory per second [20]. If decode reads every weight once per generated token, as a dense model does, then a model sized to fill that memory tops out near 4.7 tokens per second. Attention over the context, sampling and framework overhead take their cut from there. Use half the memory and the ceiling roughly doubles.

Apple puts UltraFusion die-to-die bandwidth above 4.4 TB/s [3], roughly 3.7 times the external memory figure [19]. On a quad-die part assembled from two dual-die M5 Max SOCs [2], that ratio says DRAM is the limit during generation, not the seam between dies. In a memory-bound decode, the rest of the silicon waits: the 80 GPU cores each carry a neural accelerator, and there is a 32-core neural engine alongside them [10].

The review's headline says the machine "outpaces DGX Spark and Threadripper" [9], and its verdict lists "1.2 TB/s memory throughput is excellent for local AI models" as a pro [8]. The published figures are a Geekbench 7 win [12] and Cinebench clock estimates of 4.6 GHz single-core and 4.3 GHz multi-core [11]. Tom's Hardware says it compared the system against workstations instead of desktops because it does not test many workstations [17]. For a local-serving comparison to move to your stack, the other box has to run the same quantization, the same context length and the same runtime. A bandwidth ratio does not survive a switch from a 4-bit dense model to a mixture-of-experts with few active parameters. The review does not price the machine against rented GPU hours, so the utilisation maths stays with the buyer.

Sustained inference in this enclosure is a thermal question. The chassis is 7.7 by 7.7 inches and 3.7 inches tall [24]. The M5 Ultra version weighs 8 pounds, two more than the M5 Max variant, because Apple swaps the thin aluminium stack and copper heat pipe for a copper fin stack and a copper vapour chamber [15]. Two pounds of copper is the cost of holding those clocks in a box that size.

Apple discontinued the Mac Pro earlier this year, leaving the Mac Studio as its top-end desktop [13], and this is the only machine Apple currently puts the M5 Ultra inside [25]. There is a Kensington lock slot, on the bottom, which needs a special adapter to use [16].

What to watch

  • Whether the promised 512GB memory option ships, and what Apple charges for it, since capacity cannot be added later.
  • Published tokens-per-second figures for the same quantized model on M5 Ultra, DGX Spark and Threadripper under the same runtime.
  • Whether Apple puts the M5 Ultra in a second machine, which would give buyers a form factor choice at that memory tier.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories