Product1 publisher3 min readPublished
Huawei's 2035 token forecast assigns more than 90% of the traffic to agents
The company published the number days before its Connect 2026 show, alongside ten directions that name the cluster hardware, memory systems and chip method it sells. The agent share is the part a budget has to absorb.
The Product Desk · Product desk

What happened
- Huawei published two reports days before Huawei Connect 2026 opened: Intelligent World 2035, and a Global Digitalization and Intelligence Index built with the Institute of Economics at Tsinghua University.
- The forecast underneath the rest expects global annual token consumption to grow 100,000-fold by 2035, with agents generating more than 90% of that traffic.
- The companion index forecasts more than $27trn in cumulative AI economic value over five years and 19.14% compound annual growth in digital infrastructure investment, passing $4trn by 2030.
Compiled by The Product DeskSomething wrong?How this is made
Why it matters
- decision A budget built on prompted turns cannot price the workload Huawei describes, so teams sizing inference for next year have to model turns no user initiated and tool calls they cannot count in advance.
- cost Huawei's own two targets do not net out to a cheaper bill: 100,000-fold volume against a 1,000-fold cheaper task still leaves spending up 100-fold, and the customer pays that.
- exposure Huawei sells the networking, computing and storage the forecast requires, so each figure in it also sizes the market for its own factories.
- precedent With an agent share now in print from one supplier, buyers have a number to put to the others, and a reason to ask what population it was measured on.
A team that has shipped a chatbot knows what a turn costs: one prompt, one answer, a bill you can multiply by monthly actives. The workload in Huawei's report does not behave that way. In its definition a chatbot responds when prompted, while an agent perceives, reasons, plans, calls tools and keeps learning [6]. One human request becomes an unknown number of model turns, and every token costs power, memory bandwidth and network capacity [7].
Take the forecast at face value and the growth rate is easier to hold in your head than the multiple. Measured from the 2026 publication year, 100,000-fold by 2035 is nine years of compounding at about 3.6 times a year [22]. The agent share carries nearly all of it. More than 90% of that total works out to at least 90,000 times the volume running today [23].
Each of the ten directions maps onto a product Huawei sells. Huawei wants computing clusters to scale 100-fold and the cost of an agent task to fall 1,000-fold, and says the industry has to move to SuperPoD-style systems to get there, which is what it spent Connect 2026 selling [8][9]. The storage direction asks for memory with causal accuracy and traceable provenance, the case for the context memory storage cluster it launched in Shanghai [10]. Another is the Tau Scaling Law, the chip design method it unveiled in May, which proposes time scaling in place of geometric scaling [11]. That method matters to a firm that cannot buy the newest lithography. An Agent OS direction is summarised as coordination, execution, memory and connectivity multiplied together, then raised to the power of evolution [21].
Those first two targets sit oddly together for anyone doing the budget. Volume up 100,000-fold against a unit cost down 1,000-fold still leaves spending up 100-fold, if tokens per task hold steady [24].
"Visions are painted in words, but measured in deeds," David Wang wrote in the foreword [13]. "Agentic AI is a key variable in this transformation," he wrote [12]. The Next Web, which read both reports, noted that none of it is disinterested, since a 100,000-fold increase in token consumption requires an enormous amount of new infrastructure and Huawei manufactures that infrastructure [18]. The same publication applied the caution to American projections too: "We have already gone through those numbers and found the methodologies thin" [19].
For a team sizing inference spend next year, the vendor's multiple is the wrong number to plan against. Two you can measure inside your own product are tokens per completed task, and the share of model turns no user initiated. The first tells you what an agent costs when it works. The second tells you how fast that cost grows with no change in signups, and it is the one Huawei puts above 90% by 2035 [5]. A Chinese state report in September described the same move, from models to agents, with the compute burden going from training to inference [17].
What to watch
- A base-year token volume from Huawei, which would let the 100,000-fold multiple be checked against real traffic.
- Whether SuperPoD deployments get anywhere near the 1,000-fold fall in agent task cost, and on what task definition.
- Whether American vendors put an agent share of token traffic in print with the method behind it.