Published · 17h agoLeadership3 min read
Nvidia's agent-workload lead scales with interactivity, and the second-source budget line does not
SemiAnalysis's AgentX replays recorded coding-agent sessions instead of fixed prompts. On that traffic, the reported Nvidia-AMD cost gap grows as the interactivity target rises.
Context for builders, not their beat.See today for builders
What happened
- SemiAnalysis released AgentX on August 24, an open source benchmark that replays recorded coding-agent sessions against production inference stacks instead of fixed-length prompts.
- On GLM 5.3 through open source SGLang, it reports Nvidia hardware up to five times better on cost efficiency than AMD at 150 output tokens per second per user.
- The replayed traffic has a median input of 142,000 tokens against a median output of 444, with 3.84 seconds of idle time between turns while the agent waits on a tool.
- The corpus came from a proxy on SemiAnalysis staff using Claude Code and Codex: more than 8,000 sessions and 610 billion tokens.
- AMD's ATOM engine beats a GB300 NVL72 rack running vLLM on price-performance across part of the Kimi K3 curve between 40 and 60 seconds of latency.
Compiled by The Board RoomSomething wrong?How this is made
Why it matters
- exposureSemiAnalysis puts the floor below free: at that operating point, accelerators donated at zero cost would still lose on cost per token once hosting and power are paid, which leaves whoever signed...
Claim ledger
Ranked by verification strength, evidence, and original report placement.
- [1]
SemiAnalysis published AgentX on August 24, an open source benchmark that replays real coding-agent sessions against production inference stacks rather than the fixed-length prompts most comparisons use.
ReportedView cited source - [2]
On GLM 5.3 served through open source SGLang, SemiAnalysis reports Nvidia hardware reaching up to five times better cost efficiency than AMD at 150 output tokens per second per user.
ReportedView cited source - [3]
SemiAnalysis says that at that operating point, competing accelerators handed over at zero cost would still produce a higher cost per token once hosting and power are counted.
ReportedView cited source - [4]
The published B200 comparison against MI355X on the same model shows the advantage moving with interactivity: near 57% at 108 output tokens per second per user and 247% at 141.
ReportedView cited source - [5]
Forbes cautions that the five-times figure should be read as a claim about one configuration, not a market-wide ratio.
- [6]
SemiAnalysis argues that long-context multi-turn agentic sessions now dominate production inference traffic.
Sources & coverage · 1 publisher
The reporting this story was synthesized from, earliest first. Every link goes to the original.
- forbes.comJanakiram MSV, Senior Contributor21h agoOne Agent Benchmark Puts Nvidia 5x Ahead Of AMD On Cost
Additional citations
- Forbes (Janakiram MSV)
- SemiAnalysis


