By Janakiram MSV
Publication Date: 2026-08-25 03:45:00
SemiAnalysis published AgentX on August 24, an open source benchmark that replays real coding-agent sessions against production inference stacks rather than the fixed-length prompts most comparisons use.
On GLM 5.3 served through open source SGLang, SemiAnalysis reports Nvidia hardware reaching up to five times better cost efficiency than AMD at 150 output tokens per second per user. The firm goes further at that operating point. Competing accelerators handed over at zero cost would still produce a higher cost per token, it says, once hosting and power are counted.
Read that as a claim about one configuration, not a market-wide ratio. The published B200 comparison against MI355X on the same model shows the advantage moving with interactivity, near 57% at 108 tokens per second per user and 247% at 141.
SemiAnalysis argues that long-context multi-turn agentic sessions now dominate production inference traffic. If that holds, buyers who spent two years waiting for a second source to soften Nvidia’s pricing are watching the gap widen where it counts.
What AgentX Actually Measures
AgentX is a trace replayer, not a prompt generator. SemiAnalysis disclosed that it ran a proxy intercepting Claude Code and Codex requests from its own staff, building a corpus of more than 8,000 sessions and 610 billion tokens. The version 1.0 public subset is 393 Claude Code sessions released under Apache 2.0.
Content is transformed into session-scoped chained hash blocks of 64 tokens each. This…



:max_bytes(150000):strip_icc()/GettyImages-2205219950-75903e531f684df38fd1d0f9170406fa.jpg?resize=1500,1039&ssl=1)