Invest1 publisher3 min readPublished
Meta's four MTIA generations show inference leaking from Nvidia one workload at a time
Four chip generations on a six-month cadence, a Broadcom deal running to at least 2029, and hundreds of thousands of prior-gen parts in production. Meta still calls them complements to Nvidia GPUs.
The Investor · Invest desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction
What happened
- Meta unveiled four new generations of its Meta Training and Inference Accelerator (MTIA) chip family on March 11, 2026: the MTIA 300, 400, 450 and 500.
- On April 14, 2026, Meta formalised an expanded partnership with Broadcom to co-develop MTIA chips through at least 2029.
- The MTIA lineup uses a modular chiplet design that lets Meta iterate roughly every six months, far faster than the traditional GPU development cycle.
- The earlier MTIA 100 and 200 series are already deployed in production, with hundreds of thousands of those chips running inside Meta's data centres handling ranking and recommendation workloads that power the Instagram feed and Facebook's ad-targeting engine.
- Some MTIA chips have been tested with Meta's Llama large language models.
Compiled by The InvestorSomething wrong?How this is made
Why it matters
Meta unveiled four new generations of its Meta Training and Inference Accelerator line on March 11, 2026, the MTIA 300, 400, 450 and 500 [1], and on April 14, 2026 formalised an expanded partnership with Broadcom to co-develop the family through at least 2029 [2]. The announcement matters less than the clock behind it: the modular chiplet design lets Meta iterate roughly every six months, far faster than a conventional GPU development cycle, according to Crypto Briefing's account [3].
Four generations at a six-month step is about two years of declared roadmap [14], and the Broadcom agreement extends the commitment at least three years beyond signing [15]. That is the difference between a science project and a supply chain.
The installed base is the part worth reading twice. The earlier 100 and 200 series are already deployed in production, with hundreds of thousands of chips inside Meta's data centres handling ranking and recommendation, the work behind the Instagram feed and Facebook's ad targeting [4]. Some have been tested against Meta's Llama models [5]. The newer 450 and 500 variants are aimed further into generative AI inference, with improved high-bandwidth memory to feed transformer throughput [6]. An internal memo surfaced in July 2026, per the same report, said production of the latest part, code-named Iris, would begin in September 2026 after six weeks of successful testing [7].
None of this reads as a rupture with Nvidia, and Meta does not present it as one. The company continues significant GPU purchases from both Nvidia and AMD, and positions MTIA as a complement rather than a replacement, targeting inference workloads where specialised silicon can beat a general-purpose GPU on performance per watt [9]. That framing is accurate and also the whole point. Ranking and recommendation went first because the workload is stable, high volume, and Meta owns both ends of it. Generative inference is next because the memory bandwidth is now there [6]. Each transfer is small enough to be uncontroversial and permanent enough not to come back.
The demand backdrop is why the leakage is easy to miss in Nvidia's order book. Meta plans to scale compute capacity from 7 gigawatts in 2026 to 14 gigawatts by 2027 [8], a doubling inside a year [13]. Against that, absorbing an ever larger share of inference in-house is compatible with buying more GPUs than last year, not fewer. The metric to distrust is unit shipments. The metric that moves is which workloads sit on which silicon.
Meta is late to this rather than early. Google's TPUs are in their sixth generation, Amazon has Trainium and Inferentia for AWS customers, and Microsoft has Maia [10]. Nvidia's defence remains CUDA, the proprietary software ecosystem that makes switching hardware painful [11], and CUDA is a stronger lock on heterogeneous research workloads than on a recommendation service one company writes, owns and reruns billions of times a day.
Watch whether MTIA moves to training rather than inference only, which the report identifies as the far more direct challenge and the place where the MTIA 500's memory bandwidth would matter [12]. Watch, too, whether Iris actually enters production on the September 2026 schedule [7]; a six-month cadence that slips once is a two-year roadmap that is really a four-year one.