Build1 publisherNot yet confirmed elsewhere3 min readPublished
Token-matched, agent memory modules lose to simply giving the actor more steps
A preprint from ServiceNow AI Research and university co-authors reports that a plain web agent with a longer horizon matches or beats AWM, ASI and ReasoningBank, often on fewer tokens.
The Engineer · Build desk
What happened
- The paper studies online augmentation, where module overhead is paid on every task, and re-evaluates the benefits of memory, workflow and skill modules under a fixed total inference budget.
- The authors compare AWM (Agent Workflow Memory), ASI (Agent Skill Induction) and ReasoningBank against a token-matched vanilla baseline that uses the same budget for additional actor steps.
- Across three WebArena domains and three models, Gemini 3 Flash, GPT-5.4-mini and Qwen 3.6-27B, the vanilla baseline matches or surpasses all three augmentation methods in aggregate success rate while often using fewer total tokens.
- The authors observe a similar trend on WorkArena-L1 with Qwen 3.6-27B, indicating the effect extends to enterprise knowledge-work tasks.
- The augmented methods use the default 10-step actor horizon, while the budget-matched baseline, called Vanilla-Increased-Budget (Vanilla-IB), allocates the same inference budget to a longer interaction horizon.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
Researchers at ServiceNow AI Research, with co-authors at ETS Montreal, the University of British Columbia and McGill, re-ran three popular online augmentation methods for web agents with the auxiliary token spend counted against the same budget as the actor [12][1]. Under that accounting, a vanilla actor that spends the identical budget on extra observe-and-act steps matched or surpassed all three in aggregate success rate, and often used fewer total tokens [3].
The comparison covers Agent Workflow Memory, Agent Skill Induction and ReasoningBank against what the authors call Vanilla-Increased-Budget: the augmented systems keep the default 10-step actor horizon, while the baseline converts the module budget into a longer interaction horizon [2][5]. The result holds across three WebArena domains and three models, reported as Gemini 3 Flash, GPT-5.4-mini and Qwen 3.6-27B, and the authors observe the same trend on WorkArena-L1 with Qwen 3.6-27B, which they read as evidence that the effect reaches enterprise knowledge-work tasks rather than just shopping and forum benchmarks [3][4].
The mechanism is not mysterious, and that is the point. The paper frames augmentation as an allocation problem: tokens spent inducing, retrieving, verifying or injecting reusable knowledge are tokens not spent looking at the current page, reasoning about its state, or taking one more action [7]. In the online setting the authors study, where the agent updates itself from trajectories collected during the evaluation run rather than from a skill library built beforehand, that overhead is paid on every task [1][15]. The reason this has gone unnoticed is a reporting habit: module token cost is rarely printed next to the actor's inference cost, which makes it impossible to tell whether a reported gain fits inside a fixed budget or simply bought more compute [6].
Two secondary findings matter more to anyone running these systems than the headline. First, online augmentation ties performance to task order, because accumulated memory depends on the sequence in which tasks arrive rather than the full task distribution, and the sequential dependence between tasks limits task-level parallelism [8]. That is an operational cost on top of the token cost: you lose the ability to fan work out. Second, the authors report that run-to-run variance materially affects outcomes and argue it should be treated as a core evaluation criterion for online web agents [9]. Most of the published deltas in this area are small enough that this is a substantive objection.
The authors do not claim the modules are useless. Their stated conclusion is that skills, defined here as executable code such as a Python function wrapping a reusable sequence of low-level actions, and workflow memory can help in specific domains, but that the apparent gains often disappear against a budget-matched actor [10][13]. They also offer a reason the older results looked better: prior work largely assumed base models could not finish tasks end to end, and more recent frontier models may need less external scaffolding [11].
Worth noting what the supplied excerpt does not contain: the comparisons are stated directionally, without the numeric success rates or token counts behind them [16]. Read the tables before acting on this. The practical test is cheap either way, which is the uncomfortable part: rerun your own agent with the memory stack off and the step limit raised to the same total token spend, and log both numbers side by side.
Clarity's read
What the record supports and how the coverage leans. The claims behind it follow.
Reality
- Evidence40
- Adoption
- Insufficient
- Hype gap+25
- Incentives45
- Confidence42
Claim ledger
Ranked by verification strength, evidence, and original report placement.
- [1]
The paper studies online augmentation, where module overhead is paid on every task, and re-evaluates the benefits of memory, workflow and skill modules under a fixed total inference budget.
- [2]
The authors compare AWM (Agent Workflow Memory), ASI (Agent Skill Induction) and ReasoningBank against a token-matched vanilla baseline that uses the same budget for additional actor steps.
- [3]
Across three WebArena domains and three models, Gemini 3 Flash, GPT-5.4-mini and Qwen 3.6-27B, the vanilla baseline matches or surpasses all three augmentation methods in aggregate success rate while often using fewer total tokens.
- [4]
The authors observe a similar trend on WorkArena-L1 with Qwen 3.6-27B, indicating the effect extends to enterprise knowledge-work tasks.
- [5]
The augmented methods use the default 10-step actor horizon, while the budget-matched baseline, called Vanilla-Increased-Budget (Vanilla-IB), allocates the same inference budget to a longer interaction horizon.
- [6]
The token cost of auxiliary modules is rarely reported next to the actor's inference cost, which makes it hard to tell whether observed gains are justified within a fixed budget.
- [7]
The paper frames the issue as an allocation problem: tokens spent inducing, retrieving, verifying or injecting reusable knowledge are tokens unavailable for direct task execution, such as observing the current page, reasoning over its state, or taking additional actions.
- [8]
In the online setting, auxiliary costs are incurred repeatedly, the knowledge accumulated by memory or skill modules depends on the order in which tasks are encountered rather than the full task distribution, and task-level parallelism is constrained by the sequential dependence between tasks.
- [9]
The authors report that run-to-run variance materially affects outcomes and should be reported as a core evaluation criterion for online web agents.
- [10]
Following Agent Skill Induction and SkillWeaver, the paper uses 'skill' to mean executable code, typically a Python function, that wraps a reusable sequence of low-level actions.
- [11]
Prior work on augmentation methods largely assumed base models lacked the capability to complete tasks end-to-end; the authors note that more recent frontier models may no longer face that constraint and may require less external scaffolding.
- [12]
The paper lists affiliations of ServiceNow AI Research, ETS Montreal, the University of British Columbia and McGill University.
- [13]
The authors conclude that skills and workflow memory can be useful in specific domains, but their apparent gains often vanish against a budget-matched actor.
- [14]
The bottom panel of Figure 1 previews the main WebArena pattern: once total token usage is counted, a budget-matched vanilla agent achieves the highest average success rate while using fewer tokens than the online augmentation methods.
- [15]
The study focuses on the online setting, where agents update their behavior using trajectories collected during evaluation, rather than relying on offline skill libraries constructed before deployment.
- [16]
The supplied excerpt states the comparative outcomes qualitatively (matches or surpasses in aggregate success rate, often fewer total tokens) and reports no numeric success rates or token counts.
Sources
1 independent publisher whose own reporting we read for this story.
Topics and entities
Follow any of these and your For You feed starts watching them — no settings page required.
Topics
- Web AgentsFollow
- Inference Budget AllocationFollow
- Agent Evaluation MethodologyFollow
- Agent Memory and Skill ModulesFollow
Entities
- ServiceNow AI ResearchFollow
- ÉTS MontrealFollow
- University of British ColumbiaFollow
- McGill UniversityFollow
- Agent Workflow MemoryFollow
- Agent Skill InductionFollow
- ReasoningBankFollow
- SkillWeaverFollow
- Vanilla-Increased-BudgetFollow
- WebArenaFollow
- WorkArena-L1Follow
- Gemini-3-FlashFollow
- GPT-5.4-miniFollow
- Qwen 3.6-27BFollow