Build1 distinct publisher3 min readPublished
CRMArena-Pro reports about 58% single-turn success and 35% multi-turn over nineteen expert-validated tasks, and the per-skill breakdown inside those averages is the part that should decide where an agent gets pointed.
The Engineer · Build desk
Compiled by The EngineerSomething wrong?How this is made
The multi-turn figure comes off the same benchmark, with a persona-guided counterpart in the loop instead of the problem handed over whole [3]. The 23-point gap is bought entirely by that interaction layer [1]. In relative terms it is worse than the subtraction suggests: 23 of 58 points is roughly 40% of the single-turn result gone [2], and 35% means about two of every three multi-turn attempts end somewhere other than done [3].
The abstract does not decompose those failures, so anything about the mechanism is inference. Two candidates fit the design: getting the missing field out of a counterpart who answers in character, and holding earlier constraints while the conversation moves. The authors' own remedy list puts multi-turn reasoning first [10], which is consistent with either.
Workflow Execution runs the other way, above 83% single-turn, which is 25 points above the aggregate [4] [5]. Before that number moves any planning it has to survive a transfer test. It was measured inside the sandbox environment and data generation pipeline inherited from CRMArena [6], on data synthesized for the purpose [7]. Your org has fifteen years of half-migrated custom objects and a required field nobody fills in. The abstract credits the results to "leading LLM agents" without naming them in the text supplied here [5], so it is also a benchmark against whichever model you had already decided to buy.
Taken at face value, 83% per step still does not chain. One step in six fails [6]. Three steps run unattended, assuming the errors are independent, gives 0.83 x 0.83 x 0.83, or about 57% [4], which lands back at the aggregate single-turn number [2]. That is the actual design constraint. The unit of automation is one step that produces an artifact a person can see, not a sequence. If you want the sequence, you put a checkpoint between links, and the checkpoint exists to catch the 17%.
In my context, back-office CRM work with a queue and an audit trail, that scopes cleanly enough: the agent touches the system of record, the person touches the customer. The nineteen tasks span sales, service, and configure-price-quote work in both B2B and B2C variants [1], and I would want to know which of them sit in the Workflow Execution bucket before committing to more than that. The paper's closing ask is a research agenda, better multi-turn reasoning, confidentiality adherence, and versatile skill acquisition [10], and those arrive on somebody else's schedule. The scoping decision gets made with 35% as the operating assumption.
This is revisable. A per-task table showing which tasks cleared 83%, and with which models, would let me widen the scope or narrow it further. The aggregate, on its own, only supports the narrow read.
Ranked by verification strength, evidence, and original report placement.
The abstract attributes the reported figures to "leading LLM agents" and does not name specific models in the supplied text.
CRMArena-Pro is an expert-validated benchmark with nineteen tasks across customer sales, service, and configure, price and quote (CPQ) processes, for both Business-to-Business and Business-to-Customer scenarios.
Experiments show leading LLM agents achieve approximately 58% single-turn success rate on CRMArena-Pro.
Performance drops significantly in multi-turn settings, to 35%; the benchmark incorporates multi-turn interactions guided by diverse personas.
Among the business skills evaluated, Workflow Execution is notably more tractable, with top-performing agents surpassing 83% success rate in single-turn tasks, while other skills present greater challenges.
CRMArena-Pro builds upon the sandbox environment and data generation pipeline of the earlier CRMArena benchmark.
Distinct publishers with included, body-backed reporting in this cluster.
1 article · August 31, 2026
Follow any of these and your For You feed starts watching them — no settings page required.
invest
Google's 180-config sweep: extra agents cut sequential-task scores by 39 to 70%1 distinct publisher
invest
Salesforce's double digits, minus Informatica: agentic AI is real and still 2% of revenue1 distinct publisher
build
Notion's agent stack is live, not slideware, and it only changes one of your decisions1 distinct publisher
invest
Google says frontier models already know the facts they get wrong. That is a budget decision.1 distinct publisher
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Primary and unusually well documented, entirely self-reported
This is the artifact itself rather than a report about it, and the methodology is laid out to a level most benchmark write-ups skip: 25 interconnected Salesforce objects, 21 latent variables, 29,101 and 54,569 records for the B2B and B2C orgs, plus expert studies with CRM professionals attesting realism. What holds the score down is that nothing here has been checked by anyone outside the authors, and the available text never breaks the aggregate into named models, so 'leading LLM agents' cannot be audited.
Scores exist; uptake is unobserved
A published set of results is not the same as anyone using the benchmark. We have no evidence in this reporting of a third party running CRMArena-Pro, no leaderboard submissions, no vendor adopting it in evaluation practice, and no deployment of an agent whose behaviour these numbers governed. Rating adoption from a paper's existence would be inventing the part that matters.
The framing runs cooler than the findings
Papers usually oversell; this one writes 'solely 58%' about its own results and closes on a capability gap, so the pressure runs the other way. If anything the single aggregate number understates what the per-skill spread shows — an 83% workflow ceiling against everything else — which is the detail an implementer actually needs. The one unchecked flourish is the claim to be the first benchmark designed for this breadth of business scenarios, a priority claim no one outside the authors has tested.
The authors set the exam and marked it
CRMArena-Pro is the same team's extension of CRMArena, and a benchmark that reports low scores is a benchmark that justifies its own existence — a difficulty this steep is exactly what makes the work publishable. Against that, the incentives here cut across vendor interests rather than with them: naming gemini-2.5-pro as the strong performer and reporting near-zero confidentiality awareness helps no product story, and there is no pricing, licensing or commercial claim anywhere in the text.
Firm on what is stated, thin on what is checked
We can be confident about what the paper says, since we are reading the paper: the figures are consistent between abstract and introduction, and the arithmetic on them is straightforward. Confidence in the figures themselves is lower — a first-version preprint, one publisher, no replication, and an unnamed model set behind the averages. The derived compounding numbers assume independent step failures, which real agent trajectories rarely honour.