Invest1 publisher3 min readPublished
Microsoft's log-reading CASD method cuts agent prompt tuning to about $1.60 a pass
Microsoft researchers say CASD writes better agent prompts from saved logs for about $1.60 a pass, over 22 times cheaper than iterative tuning. Outside one telecom test its accuracy edge over GEPA is thin, so price and setup carry the case.
The Investor · Invest desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened
- A coding agent reads the full corpus of saved agent trajectories in one pass, extracts recurring failure modes and representative episodes, and writes behavioural rules into a new system prompt.
- Across the tested tasks, the optimized prompts averaged a gain of 16.6 percentage points over the unoptimized baseline, according to the paper.
- The paper says CASD kept its lead even when rival methods were given extra validation data and unrestricted environment access.
- GEPA, the main comparison, is described in the paper as widely adopted at Databricks, Shopify, Dropbox and OpenAI, among others.
Compiled by The InvestorSomething wrong?How this is made
Why it matters
- cost For agent teams, a larger saving than the compute bill is staff time: CASD drops the held-out validation set that iterative tuners need someone to build and keep current.
- exposure Tuning without live environment access stops an optimizer from testing draft prompts on systems where the agent can act on customer accounts or production tools.
- constraint Output quality is capped by the saved logs, so a team whose trajectories are narrow or unrepresentative can end up with rules that do not generalize.
On the paper's own ratio, a validation-gated iterative pass costs more than $35.20, because CASD's roughly $1.60 is put at over 22 times cheaper [1][2][1]. The direct saving is at least about $33.60 a pass [2]. For a single tuning job, $33.60 is a small sum. It becomes a budget line only for a team re-tuning many agents often, and the source does not say how often teams re-run optimization or how many passes an iterative run takes.
The $1.60 is also a marginal price. CASD works from saved trajectories and needs no iterative search or live interaction with the environment [3]. Those past episodes were paid for when the agent first ran them. A team that has not kept its agent logs has nothing to feed the method.
Against GEPA, the margins are uneven. CASD led by 9.3 points on ALFWorld, 0.8 on tau2-bench retail and 21.7 on tau2-bench telecom, and trailed by 9.4 on SpreadsheetBench-Verified, where GEPA scored 60.7 to CASD's 51.3 [3][11]. Averaged across the four, CASD's lead is 5.6 points [4]. Take telecom out and the average edge across the other three is about 0.2 points [5]. Telecom supplies 21.7 of the 22.4 net points CASD gained across the four tests [6].
The write-up puts the gains down to scope. Iterative optimizers typically look at slices of data at a time, and patterns that only appear across many episodes are harder to spot that way [19].
Microsoft is also arguing against its own earlier work. SkillOpt, a validation-gated iterative method that treats agent skills as trainable text parameters, came out of Microsoft in June 2026 [15]. The new paper's six authors, Sumit Gulwani among them, report CASD beating SkillOpt-style search across its metrics [14][6]. Cryptobriefing reads the result as Microsoft saying that, in many cases, its earlier approach is more work than necessary [20].
Three outcomes look plausible. Results that hold on production logs would leave iterative loops a role mainly on structured work like spreadsheets. A corpus of poor logs would move the effort from building validation sets to curating trajectories. An outside replication that narrows the telecom gap [7] would leave the GEPA comparison close to a tie on accuracy, with CASD ahead on price and setup.
I think the third is the likeliest. It still favours CASD for teams that already store large volumes of agent logs, because a 0.2-point average edge at a 22nd of the cost is a good deal [5][2]. The counter-case is the spreadsheet benchmark, where GEPA's 9.4-point lead shows an iterative loop still paying on structured tasks [11]. The view is wrong if an outside replication on production logs finds CASD trailing GEPA across most task types, not only spreadsheets.
What to watch
- An independent replication of the CASD versus GEPA results, especially the 21.7-point telecom gap, on agent logs outside the paper's benchmarks.
- Whether GEPA users named in the paper, such as Databricks or Shopify, publish comparisons run on their own production trajectories.
- Cost figures for CASD measured on real trajectory corpora, including how large a log store the method needs to beat iterative tuning.