Skip to content

Build1 publisher3 min readPublished

Fifty hops, one budget: why more GPU capacity won't fix agent latency

Akamai says half of enterprise AI deployments miss their own latency targets at peak load. The arithmetic points at where the work runs, not at how much GPU sits behind it.

The Engineer · Build desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened

  • Akamai's State of AI Inference 2026 report surveyed 200 AI practitioners.
  • The report found that 50% of enterprise AI deployments are failing to meet their latency demands at peak load.
  • 82% of organizations say their most critical use cases require end-to-end response times of 500 milliseconds or less.
  • 64% of organizations now require end-to-end response times of less than 250 milliseconds for their most important use cases.
  • When an agent built on a framework like LangChain, CrewAI or Pydantic AI receives a user request, it can fan out into dozens of sequential operations such as a reasoning call, a tool invocation, an API lookup or a context retrieval, then a further reasoning call to decide what to do with what came back.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

Akamai's State of AI Inference 2026 report, a survey of 200 AI practitioners, found that half of enterprise AI deployments are missing their own latency targets at peak load [1][2]. The number should bother operators because the targets are not decoration: 82% of organizations said their most critical use cases need end-to-end responses in 500ms or less, and 64% said under 250ms [3][4].

An agentic request is not one inference call. A request arriving at an agent built on LangChain, CrewAI or Pydantic AI can fan out into dozens of sequential operations, including a reasoning call, a tool invocation, an API lookup, a context retrieval, then another reasoning call to decide what to do with what came back [5]. Akamai's post argues that every hop crossing a wide-area network to reach a centralized data center adds transport time, and that a chain of 50 hops can multiply that into seconds on its own, whatever the token generation rate [6].

Do the division. A 500ms budget spread across a 50-hop chain leaves 10ms per hop, inference included [1]. At the 250ms target, 5ms [2]. A larger GPU allocation does not return any of that.

The supporting evidence is the part worth keeping. A paper posted to arXiv in November 2025 found that CPU-side processing can account for up to 90.6% of total latency in agentic workloads [7]. If that holds, everything else, token generation included, is at most 9.4% of the wall clock [3], so making the model twice as fast buys you under 5% [4]. That is the sentence to read out in the next capacity meeting. As the Akamai piece puts it, "you can't brute-force your way out of a wait state" [8].

Teams find out late because the instruments are wrong. Most LLM-serving benchmarks measure tokens per second and GPU utilization on a single box, which is the right test for one model answering one prompt and the wrong one for a 50-hop response that crosses a WAN four times to reach four services [9]. In the author's words, staging may pass because it tests the model, while "production tests the whole chain, including every hop your serving engine was never designed to see" [10]. The observable symptom is GPU idle time: the reasoning step finishes in a few hundred milliseconds, then waits on tool calls to CPUs in distant data centers [11].

The volume is arriving regardless. LangChain's State of Agent Engineering 2026 survey of more than 1,300 professionals found 57.3% of organizations running agents in production, up from 51% a year earlier [12], a rise of 6.3 percentage points [5], with latency now the second-most-cited barrier to production behind output quality [13].

A note on provenance: the article is written from inside Akamai, whose commercial answer is distributed placement, and it draws the analogy to the company's own founding by MIT researchers Tom Leighton and Danny Lewin in answer to a challenge from Tim Berners-Lee [14]. Ari Weil, who leads product marketing for Akamai's cloud computing business and ran the research, says "the enterprise AI honeymoon phase is over... they are hitting the latency wall" [15]. The self-interest is visible. The division is still the division.

Watch for anyone publishing a benchmark that measures a whole chain rather than a single box [9], because until one exists, procurement will keep converting a placement problem into a GPU purchase order. Watch chain length too: at the same 500ms target, a 100-hop workflow cuts the per-hop budget to 5ms [6]. And check your traces for GPU idle time during tool calls before signing anything [11].

Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories