Invest1 distinct publisher2 min readUpdated
The best of 12 models passed 65.36% of 507 business tasks on the first attempt and only 25.25% across all 20 trials. Agent procurement should be priced on the second number.
The Investor · Invest desk
Compiled by The InvestorSomething wrong?How this is made
Independence would have made this far worse. If a 65.36% first-try rate [6] were an honest coin flip on every attempt, clearing all 20 trials of a task would happen about 0.02% of the time [1]. Microsoft reports 25.25% [7], roughly 1,250 times higher [2]. That discrepancy is the useful part of the release, because it says the failures are not noise sprayed evenly across the task set. They are concentrated.
Read the two figures together and the shape emerges. About a quarter of the 507 tasks [5] the leading model does dependably. The other 40.11 points of its first-pass score [3] come from tasks it fails at least once in 20 attempts, which is 61% of all its first-try wins [4]. So on this task mix, a single successful run is more likely than not to be a sample from the band the model cannot reproduce. That is what a vendor demo is: one trial, self-selected, with the harness supplied by the seller.
The cost of the gap falls on operations, not on the model line item. At 65.36%, a thousand single-pass task runs leave roughly 346 wrong outcomes behind [5], and the reconciliation work scales with transaction volume rather than with whatever the next checkpoint improves. Microsoft's own framing of the alternative is instructive: transcript analysis, the standard method [4], reads the agent's account of itself, while ThinkingBox runs executable assertions against the database state after each isolated trial [3].
Measuring this properly is not free either. The benchmark design implies 10,140 trials per model and 121,680 across the 12 tested [6]. A buyer replicating the method on its own task list pays 20 times the compute of a one-pass pilot [7], which is an unusual line to defend in a budget but at least it is estimable in advance, unlike a pilot whose result is a single draw.
The caveats are real and they run one direction. Every number here is Microsoft's, produced on Microsoft's harness, published on a Microsoft blog by a Microsoft engineer [2], as part of a stated push into agent observability and governance [10]. The account does not say which of the 12 models scored 65.36% and 25.25% [11], so nobody can shop on it directly. What it does supply is a unit of measure and the code to compute it [9], and the useful contract term that follows is a pass^k floor on the customer's own tasks, verified against the system of record, with k written into the acceptance test rather than left to the vendor.
Ranked by verification strength, evidence, and original report placement.
Microsoft has released ThinkingBox, an open-source sandbox framework designed to test whether AI agents can be trusted to handle real business tasks.
ThinkingBox was detailed in a Microsoft Command Line blog post on August 19, 2026, by Principal Machine Learning Engineer Liang-Chun Tsai.
ThinkingBox verifies back-end database records after each task, creating a clean isolated environment for every trial and running executable assertions against the final state of the system.
Traditional methods for evaluating AI agents rely on transcript analysis, reading the agent's work rather than checking the resulting system state.
Microsoft tested 12 proprietary and open-weight models across 507 tasks spanning five business domains, with each task run through 20 separate trials, using the paired benchmark ThinkingBox-Bench.
The best-performing model achieved a 65.36% pass@1 rate, completing a task correctly on its first attempt.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Precise figures, single unverified retelling
The quantitative core is specific and internally consistent - 12 models, 507 tasks, five domains, 20 trials, 65.36% pass@1, 25.25% pass^20 - and the harness mechanism is described concretely. But the cluster contains exactly one source, a crypto/tech aggregator paraphrasing a Microsoft blog post it does not link, the top-scoring model is never named, and no per-task distribution or domain breakdown is given. That is enough to report the finding, not enough to audit it.
Public release, vendor-only usage so far
There is one concrete adoption fact: the framework and benchmark are publicly released on GitHub, and the vendor itself ran a large evaluation with them. Beyond that the cluster reports no external users, no third-party benchmark runs, no downstream product integration, no download or star counts and no customer deployments, so observed uptake is confined to the publisher of the tool.
Findings sober, framing ahead of proof
The central numbers cut against agent hype rather than feeding it - a vendor publishing that its strongest tested model is consistent on only a quarter of tasks is a deflationary claim, which pulls the gap toward zero or below. What pushes it modestly positive is the surrounding framing: an unlinked primary post, an unnamed top model, four-significant-figure precision presented without verification, and a publisher inference that Microsoft intends the methodology to become an industry standard when no Microsoft statement of intent or any external adoption is shown.
Vendor authors the ruler and the scores
Microsoft designed the harness, wrote the benchmark, selected the 12 models, ran every trial and reported the numbers, while the cluster places the release inside its own agent observability and governance push - a clear interest in having its evaluation method treated as the default. The unnamed top model further shields competitive comparisons from scrutiny. The single publisher is an aggregator restating a vendor post, adding no independent check but also no evident stake of its own.
Directionally credible, thinly sourced
The mechanism story and the shape of the finding - one-shot success far above repeat-trial consistency, with silent partial failures - hang together and the derived arithmetic follows from the reported figures. Confidence stays low because everything traces to one secondary publisher with no primary link, no named model, no task-level data and no external adoption; a single additional source or the primary post would move this materially.
Follow any of these and your For You feed starts watching them — no settings page required.
product
Rillet's $100M reads as proof mid-market ERP is rip-and-replace, mostly at the cheap end1 distinct publisher
build
A UDP packet is now enough: IKEEXT RCE moves from patch queue to fire drill1 distinct publisher
product
Nebius funds $4.5bn of AI capacity on terms that pay lenders mostly in stock2 distinct publishers
product
The cheapest part in the car is now the one that stops the line1 distinct publisher
Distinct publishers with included, body-backed reporting in this cluster.
cryptobriefing.com
1 article · August 23, 2026