Skip to content

Product1 publisher3 min readPublished

Apple spreads the AI math across all six cores of its first 2nm iPhone chip

Every performance figure Apple has published for the A20 Pro is a peak number, and none of them tells a product team how long a local model takes on its fourth run, or how large a model fits in memory.

The Product Desk · Product desk

Photograph accompanying Apple spreads the AI math across all six cores of its first 2nm iPhone chip
Photo: apple.com

What happened

  • Apple's A20 Pro is the world's first 2-nanometer smartphone processor, built on the same manufacturing process as the M6 chip.
  • Its CPU is six cores, two super cores plus four efficiency cores, dropping the middle performance tier that the A19 Pro used, and Apple adds a neural accelerator to every core.
  • Apple says the seven-core GPU is up to 40% faster than the previous generation, and credits the new super cores with a 20% CPU performance improvement.

Compiled by The Product DeskSomething wrong?How this is made

Why it matters

  • constraint Every number available is a peak from Apple, with no sustained-throughput measurement, so a latency budget for a local model has to be built on a figure that describes the best case rather than the fourth run.
  • decision Choosing to run inference locally now means choosing which slice of the install base gets the feature, because the 2nm part is in the newest handsets and everything else is on older silicon.
  • capability With an accelerator on each of the six cores alongside 32 Neural Engine cores, model math no longer has to queue for a single block, which widens what a team can attempt while the interface is still drawing.
  • exposure If heat is the binding limit, the fix Apple is shipping is a cooling part rather than a software budget, which leaves the feature owner answering for a demo that ran fast on stage and slower in a pocket.

A phone chip that runs too hot decelerates to keep heat inside a safe range within about ninety seconds of sustained use, and phones have this problem worse than laptops, which have room for fans and are not expected to live in a hip pocket [13]. Apple's answer is vapor chamber cooling borrowed from high-performance laptops and the most recent iPad Pro: a flat vacuum-sealed metal chamber partly filled with deionized water that evaporates and recondenses to pull heat off the chip with no moving parts [12]. Apple used vapor chamber tech in the iPhone 17 as well, and the account of the new chip does not set out what changed between them [14].

The CPU core layout swaps one tier for another rather than adding to it. The A19 Pro ran two performance cores and four efficiency cores [4]. The A20 Pro runs two super cores and four efficiency cores with no performance tier between them [3], which leaves the total at six in both generations [5]. Super cores are the peak single-thread and boosted-throughput parts; efficiency cores handle low-energy and background work [19]. So Apple's claimed 20% CPU improvement, which it credits to the super cores [8], describes what happens when work lands on two of the six.

The AI-specific changes sit around that. Every core gets a neural accelerator to offload some of the math used for running models [6]. The dual 16-core Neural Engine is the same one Apple is putting in upcoming Macs [10], so 32 Neural Engine cores in a handset [11]. Memory throughput is up 50%, which Apple attributes to placing the CPU cores side by side with memory on the die to relieve bottlenecks among CPU, GPU and Neural Engine [9]. PCMag reads the whole thing as a hardware rethink built from the ground up for local AI workloads rather than an annual tweak [16], and it is the first 2nm smartphone processor, on the same process as the M6 [1][2].

A team rolling this out needs a sustained-throughput number under repeated inference and a memory capacity number, and neither is in the figures on offer. Throughput measures how fast weights move; capacity determines how many fit. Every performance claim here is Apple's own, with no independent test and no capacity figure [18].

Does the feature stay usable when the chip slows down, and does it need the newest silicon to be usable at all -- those two questions settle the decision without a device in hand. A feature that stays graceful and tolerates older silicon can go local now. One that stays graceful but only on the A20 Pro goes local on new hardware with a server path behind it. A feature that is fragile and A20 Pro-only is a demo, and one that is fragile on silicon it never needed is a server feature wearing a privacy story.

The useful exercise is to write down how many seconds your feature is allowed to take on the ninetieth-percentile phone in your install base, then look for the published figure that speaks to it. Apple's new CEO John Ternus calls the iPhone an "intelligent personal hub" [15], which points to where Apple wants the compute to sit, though it says nothing about what your fourth run costs.

What to watch

  • An independent sustained-throughput test of the A20 Pro under repeated local inference, which would show whether the vapor chamber holds the 20% and 40% figures past the first run.
  • A published memory capacity figure for the new iPhones, since 50% more throughput says nothing about how large a model fits.
  • Whether the per-core neural accelerators and the dual 16-core Neural Engine are exposed to third-party apps or reserved for Apple Intelligence features.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories