build1 distinct publisher
ServiceNow's agent benchmark tops out at 35.3% on the harder of its two enterprise tasks
AgentArch sweeps orchestration, ReAct versus function calling, memory scope and a thinking tool across 18 setups on frontier models. Because the best cell moves with the model, the grid is what you reuse and the ceiling is what you budget for.
Publishers:arxiv.org
Reality
- Evidence60
- Adoption
- Insufficient
- Hype gap+8
- Incentives55
- Confidence