Skip to content

Product2 publishers3 min readPublished

Apple's quad-die M5 Ultra puts 512GB of inference memory under the desk

Apple says the M5 Ultra reaches 512GB of unified memory at 1.2TB/s, which settles how big a local agent can get, though the September launch arrives without prices and every AI speed figure is measured against an M1.

The Product Desk · Product desk

Illustration accompanying Apple's quad-die M5 Ultra puts 512GB of inference memory under the desk

What happened

  • Apple's M6 and M5 Ultra arrive in September in updated Mac mini and Mac Studio desktops, with the M6 the company's first part on a 2nm process after the "3nm Class" label it applied across the M5 series.
  • The M5 Ultra joins two dual-die M5 Max chips over an upgraded UltraFusion interconnect to make Apple's first quad-die system-on-a-chip, reaching a 36-core CPU and an 80-core GPU.
  • Memory ceilings are the headline specification: 32GB at 170GB per second on the base M6, and 512GB at 1.2TB per second on the M5 Ultra, which Apple calls 50% more bandwidth than the M3 Ultra.
  • Thunderbolt 5 machines support RDMA clustering of up to four Mac Studio or Mac mini units over 120Gbps links, which Apple says yields up to three times faster distributed LLM inference.
  • Apple's AI numbers for the M6 are up to four times faster AI performance and up to 13.5 times faster LLM prompt processing, both measured against the M1 generation.

Compiled by The Product DeskSomething wrong?How this is made

Why it matters

  • capability A model whose weights and context previously needed rented accelerators can sit resident on one desktop, which makes duty cycle rather than peak speed the thing that justifies owning the hardware.
  • constraint Anything that exceeds one box's memory has to cross a Thunderbolt fabric roughly 290 times slower than the interconnect inside a single M5 Ultra, so the workload must shard cleanly or the pooling gain largely disappears.
  • decision Teams can start sizing model choice against a published memory ceiling now, but cannot yet set cost per hour of owned silicon against cost per token, because no price has been published.
  • exposure Whoever writes the purchase case is quoting first-party multiples against a five-generation-old baseline, and will have no independent measurement to point at when the review committee asks.

Somebody at your firm wants an agent triaging tickets overnight, on a machine nobody logs into. The number that decides whether that runs on a desk or on somebody's API is memory, not cores. Apple has now printed it: 512GB on the M5 Ultra [11], sixteen times the 32GB on the base M6 [15]. The working set either fits under that ceiling or it does not, and the 80-core GPU does not move the line [9].

The clustering claim is the part worth doing the arithmetic on. Thunderbolt 5 runs at 120Gbps, which is 15GB per second [12]. UltraFusion moves more than 4.4TB per second between dies inside a single M5 Ultra [8]. The fabric between machines is therefore about 293 times slower than the fabric inside one machine [16], and roughly 80 times slower than that machine's own memory bandwidth [21]. Four maxed Mac Studios would present 2TB of pooled memory [18], and Apple's figure for four networked desktops is up to three times faster distributed inference, not four [12]. That missing fourth is the price of the seam, stated by the vendor.

One more derived number, because the comparison Apple chose is with itself: if 1.2TB/s is 50% up on M3 Ultra [11], the older chip was at 800GB/s [17].

What the announcement does not contain is a price [22]. That is the whole problem with treating this as a buy-versus-rent reset. Rent is billed per token as consumed; buy is billed once, whole, in advance, and nobody can compare the two with one side blank. The performance claims that are available are Apple's own and baselined on M1: up to 4x AI performance, up to 13.5x LLM prompt processing [6]. M6 is the sixth generation of Apple silicon, so that baseline sits five generations back [19]. It is a good argument for replacing an old Mac and a weak argument against an API bill.

Teams imagine a fleet of desktops quietly absorbing agentic work, with the cloud invoice going to zero. The spec sheet supports something narrower. Two things decide it, and neither is peak throughput. First, whether the working set fits in the memory of one box [11]. Second, whether the duty cycle is continuous or bursty. Owning wins on arithmetic you can check only when the work fits and runs continuously, because the box is busy and the marginal token is free once bought. When the work fits but runs in bursts, you have bought idle silicon and should have rented. When the work does not fit and runs continuously, you are across the Thunderbolt seam [16] or into a data centre. When the work does not fit and runs in bursts, you are paying for an API call with extra steps.

There is a gate that overrides all four cells, and it is the one PCMag says Apple is aiming at: data that is not permitted to leave the building, which is the privacy-first, on-device case [13]. If that applies, the memory ceiling is your architecture, and the decision rests on capacity rather than cost per token.

Until prices appear, the defensible version of the local-agent case is the memory case rather than the speed case. If the working set fits in 512GB and the machine would be busy most hours of most days, the purchase argues itself. If it runs two hours a day, you have bought a very quiet space heater with excellent memory bandwidth.

What to watch

  • Prices and configuration tiers for a 512GB M5 Ultra Mac Studio, without which the owned-versus-rented comparison stays half blank.
  • Independent tokens-per-second measurement on a four-machine Thunderbolt 5 RDMA cluster against Apple's up-to-3x distributed inference figure.
  • Whether Apple documents which model sizes actually fit in 512GB, and how developers address the pooled memory across clustered machines.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories