Build1 distinct publisher3 min readPublished
AgentArch sweeps orchestration, ReAct versus function calling, memory scope and a thinking tool across 18 setups on frontier models. Because the best cell moves with the model, the grid is what you reuse and the ceiling is what you budget for.
The Engineer · Build desk

Compiled by The EngineerSomething wrong?How this is made
Four dimensions with two settings each is sixteen cells. The paper reports eighteen configurations [1][2][13]. So either one of those dimensions is not binary, or there are cells outside the factorial grid, and the front matter does not say which [13][15]. If you intend to reuse the harness rather than quote the headline, resolve that first, because the shape of the grid is the experiment.
The result I would act on soonest is the negative one. Across models, ReAct prompting inside multi-agent orchestration showed significant weakness, the authors report [5]. Function calling rides the provider's schema-constrained tool interface; ReAct puts the action into prose that your framework has to parse, and in a multi-agent setup that prose crosses a handoff before anything validates it. That explanation is mine, and this data does not test it. The paper's own related-work section points at earlier tool-calling benchmarks that already found weaknesses in tool selection accuracy, tool input accuracy, valid output formatting, latency and unnecessary calls [10]. That is the list I would instrument before blaming the orchestrator.
Then the ceiling. A 35.3% best score on the complex task means 64.7% of attempts did not succeed [3][11], and the simpler task landed at almost exactly double: 70.8 divided by 35.3 is 2.0 [12]. For either figure to describe your workload, you need the same definition of success, a tool inventory of similar size and shape, and tasks that run about as many steps. The supplied text gives none of those and does not name the models evaluated [15]. So read 35.3% as a measurement of two ServiceNow use cases under ServiceNow's own grader [7], and treat the shape of the finding, not the value, as the portable part. It is unusual for a vendor to publish a ceiling this low on the category it sells into.
Adopting the finding costs more than reading it. Eighteen configurations against two use cases is 36 graded runs per model [14], and the graded task set has to exist before any of them mean anything. The payoff is at the cheap end of the model list: larger models were more robust across architectures, but on the simpler task a smaller model under its best configuration matched them [6]. That is architecture search paying for itself in inference cost, on the easier class of task only.
The claim the paper makes against one-size-fits-all is narrower than it sounds, and stronger for it. Models did best under different architectures, and for a single model the best architecture changed between the two use cases [4]. That is not a statement about which framework to buy. It is a statement that your orchestration diagram is a per-model, per-workload result with a short shelf life, which is why the useful artifact here is the grid and the harness [1], run against tasks you graded yourself.
Ranked by verification strength, evidence, and original report placement.
AgentArch is a benchmark that evaluates 18 distinct agentic configurations across state-of-the-art large language models, presented in a paper attributed to ServiceNow with an accompanying GitHub repository.
The benchmark varies four architectural dimensions: orchestration strategy (single-agent versus multi-agent), agent style (function calling versus ReAct), memory management (complete versus summary), and thinking tool integration.
The highest scoring models achieved a maximum of only 35.3% success on the more complex task and 70.8% on the simpler task.
The benchmark reveals significant model-specific architectural preferences, and even between use cases models perform best on different architectures.
The authors observe significant weaknesses across models when ReAct prompting is used in multi-agent systems.
Larger models display more robust performance across architectures, but on simpler tasks smaller models can match performance under their best performing architecture.
Distinct publishers with included, body-backed reporting in this cluster.
1 article · August 31, 2026
Follow any of these and your For You feed starts watching them — no settings page required.
product
The market read Workday's backlog and skipped its agent adoption numbers1 distinct publisher
product
Half the incident clock goes to search, and telemetry tools cannot read the answer1 distinct publisher
build
Model-generated tool arguments cross the same trust boundary as an HTTP request1 distinct publisher
product
Twin1's $20M bet: the unit of enterprise AI is one employee, not the org1 distinct publisher
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Primary paper, nobody's checked it
Every figure a reader takes away — 35.3%, 70.8%, 18 configurations — comes from the abstract of ServiceNow's own paper, and our coverage contains nothing that tests it. That is strong provenance for what the authors claim and weak provenance for whether it holds: the text we have stops before methods, so the models scored and the definition of success are both absent, and the 18-versus-16 arithmetic of the grid is never reconciled.
Code posted, users unseen
A preprint and a GitHub link are the entire trail. Nobody outside ServiceNow is reported running AgentArch, no product or procurement decision cites it, and the paper discloses no usage of its own benchmark. Publishing a repository is availability, not uptake, so we leave this unscored rather than promote a code drop into traction.
Honest scores, generous scope
For a vendor benchmark this is unusually self-effacing: the lead result is that the best model fails most of the harder task, which nobody writes for marketing. The overreach is one word — comprehensive — carrying 18 setups on two undescribed enterprise tasks from a single company's workflow world. Small gap, and it lives in the scope of the claim rather than in the digits.
The automation vendor writes the rules
ServiceNow sells enterprise workflow automation and here defines what a good enterprise agent architecture is measured on — which orchestration strategies, which prompting styles, which memory scopes get a column at all. That is the interested part, and it sits in the dimension list rather than the results. Pulling the other way: the numbers embarrass the whole category, and the code is public, which is not how a promotional exercise is usually built.
Single account, one loose thread
Two things hold this down. The dimensions as listed multiply to sixteen while the paper counts eighteen, so part of the design is unaccounted for in what we can read; and there is no second account of any of it. What we would still stand behind is the shape of the finding — architecture choice is model-specific, ReAct inside multi-agent systems is fragile, enterprise completion rates are low — because the authors state those plainly and their own numbers cut against their interests.