BuildNot yet confirmed elsewhere1 publisher3 min readPublished
Pine AI's agent computer finishes fewer benchmark tasks than two rivals, by Pine's own count
Pine AI's cloud computer for agents resolved 27.4% of SaaS-Bench tasks in Pine's own test, below two rivals at 31.1% and 29.2%. Its higher step score and much lower token bill come from whole systems running different models, so the computer's own contribution has not been measured.
The Engineer · Build desk

What happened
- The service is in private beta for developers building products that need AI to act across websites, files and business software.
- On UniPat AI's SaaS-Bench v1.1, Pine reports a 78.3% checkpoint score for GPT-5.6 Luna on its computer. Opus 5 running in Claude Code scored 74.3%, and GPT-5.6 Sol running in Codex scored 71.1%.
- Model-token cost came to about $1.02 per task for Pine's system, against $26.50 and $20.50 for the two alternatives, according to Pine.
Why it matters
- contradiction Pine's system ranks first of three on steps passed and last on tasks finished, and a buyer paid on completed work will be judged by the second ranking.
- constraint No run holds the model and budget fixed while swapping the computer, so the benchmark cannot yet test Wang's claim that the environment is what holds agents back.
- cost Pine's token spend stays near $3.72 per resolved task even at its lower resolve rate, but the infrastructure Pine actually sells is left out of every cost figure.
- exposure Products built on the SDK put their users' sign-ins and business systems inside Pine's sandbox and takeover flow, so Pine's isolation layer becomes part of each developer's security boundary.
Dylan Wang's premise comes from networking. In Pine's launch post dated October 5, the co-founder recalled his time at Agora, where he came to see that audio and video quality could be improved by rebuilding the network underneath [3]. Pine Computer applies that idea to agents. It changes the environment the model works in, so the model does not have to make up for a computer designed around human eyes and hands [4]. Wang's bet is that a more capable model will not be enough on its own, and that agents also need a computer built differently [2].
A conventional computer-use agent runs a loop. It looks at a screenshot, decides what to click, acts, and looks at another screenshot [5]. Pine says its environment returns structure instead: the elements on the page, the actions available, and what changed after the last action [6]. The screen is still there. A user can watch it live, take over for a sign-in or an approval, and hand control back [7]. I'd expect a page-state diff to cost fewer tokens per step than an image. I'd also expect it to turn the question of whether a click worked into a lookup instead of a visual guess. Pine's evidence cannot confirm either point yet.
SaaS-Bench v1.1, from UniPat AI, covers 106 workflows across 23 applications [10]. Its checkpoint score counts intermediate steps passed, not tasks completed end to end [11]. Pine leads the two rivals on checkpoints by 4.0 and 7.2 points [16]. It trails them on resolved tasks by 3.7 and 1.8 points [17]. The gap between its two scores is the widest of the three systems: 50.9 points, against 43.2 and 41.9 [21]. If the resolve rate is counted over the 106 workflows, the systems finished about 29, 33 and 31 tasks [18]. With gaps of four tasks and two, I would want repeated runs before ranking anyone.
Each row in the table changes more than one variable. Pine's entry runs GPT-5.6 Luna on Pine Computer. The rivals are Opus 5 in Claude Code and GPT-5.6 Sol in Codex [10]. Runtimewire notes that the comparisons cover complete systems with different software and budgets, and do not isolate the effect of Pine's computer [14]. The closest pair shares a version number. To test Wang's thesis, a run would have to hold the model and the budget fixed and swap only the computer.
Pine reports about $1.02 in model tokens per task, against $26.50 and $20.50 for the alternatives [13]. The rivals spent about 26 and 20 times as much [19]. Pine's lower resolve rate does not close that gap. Its tokens come to about $3.72 per resolved task, while the cheaper rival's spend stays above $65 per resolved task at either rival's resolve rate [20]. The figures exclude infrastructure [14]. Infrastructure is what Pine sells. It runs the computer, browser and isolation layer, and it packages them with the agent runtime and permission handling as one cloud service [8][15].
Pine published its run data, so the scores can be checked [22]. It also labels its broader 2-5x speed claim as preliminary internal testing that varies by task [22]. I'd count both in Pine's favor.
Developers adopt it through an SDK and build the user experience around Pine's machine [8]. Each computer runs in its own sandbox, and developers keep their own keys [9]. Tasks can run in parallel, including on sites that have no API [9]. Pine has given one deployment example, an unnamed customer that uses its software to automate audits. Pine says that customer's team now handles 50% more work with the same people. The release does not explain how the increase was measured [23].
What to watch
- A SaaS-Bench run that holds GPT-5.6 Luna and the budget fixed while swapping Pine Computer for a screenshot harness.
- Pricing for the computer layer as the private beta opens, to set beside the $1.02 token figure.
- An independent rerun of Pine's published run data, with repeated runs on the roughly 30 tasks each system resolved.
Clarity's read
What the record supports and how the coverage leans. The claims behind it follow.
Reality
- Evidence35
- Adoption10
- Hype gap+40
- Incentives80
- Confidence40
Claim ledger
Ranked by verification strength, evidence, and original report placement.
- [1]
Pine AI introduced Pine Computer for developers building products that need AI to work across websites, files and business software; the product is in private beta.
- [2]
Pine AI co-founder Dylan Wang is betting that agents need a different kind of computer, not just a more capable model.
- [3]
In Pine's launch post, dated October 5th, Wang wrote that at Agora he saw how rebuilding the underlying network could improve audio and video quality.
- [4]
Pine's product applies the premise of changing the environment the model operates in, rather than asking it to compensate for a computer designed around human eyes and hands.
- [5]
Conventional computer-use agents often inspect a screenshot, decide what to click, act and inspect another screenshot.
- [6]
Pine says its browser and operating environment report page structure and changes: what elements are present, which actions are available and what changed after an action.
- [7]
Visual input remains available, and a user can watch a live screen, take control for sign-ins or approvals, then hand control back; human takeover can handle steps such as entering credentials.
- [8]
The SDK is the intended entry point. Pine runs the computer, browser and isolation layer; the developer builds the user experience around them.
- [9]
Pine says each computer runs in its own sandbox and developers retain their own keys; the product is designed to work across sites without APIs as well as files and business software, and to run tasks in parallel.
- [10]
On UniPat AI's SaaS-Bench v1.1, which covers 106 workflows across 23 applications, Pine reports a 78.3% checkpoint score for GPT-5.6 Luna running on Pine Computer, above 74.3% for Opus 5 with Claude Code and 71.1% for GPT-5.6 Sol with Codex.
- [11]
A checkpoint score measures the share of intermediate task steps passed, not the share of tasks completed end to end.
- [12]
Pine reports 27.4% of SaaS-Bench tasks resolved for its system, compared with 31.1% for Opus with Claude Code and 29.2% for GPT-5.6 Sol with Codex.
- [13]
Pine reports about $1.02 in model-token cost per task for its system, against $26.50 and $20.50 for the two alternatives.
- [14]
The cost figures exclude infrastructure and compare complete systems with different software and budgets; they do not isolate the effect of Pine's computer.
- [15]
Developers could otherwise assemble a browser, execution environment, agent runtime and permission-handling workflow themselves; Pine is selling those pieces as one cloud service.
- [16]
Pine's checkpoint score leads Opus 5 with Claude Code by 4.0 points and GPT-5.6 Sol with Codex by 7.2 points.
- [17]
Pine's resolve rate trails Opus with Claude Code by 3.7 points and GPT-5.6 Sol with Codex by 1.8 points, ranking it last of three on finished tasks after ranking first on checkpoints.
- [18]
If resolve rates are counted over the 106 workflows, the systems resolved about 29 (Pine), 33 (Opus/Claude Code) and 31 (GPT-5.6 Sol/Codex) tasks.
- [19]
The two alternatives spent about 26 and 20 times Pine's model-token cost per task.
- [20]
Pine's token cost is about $3.72 per resolved task; the cheaper alternative's $20.50 per task is above $65 per resolved task at either rival's resolve rate.
- [21]
The gap between checkpoint score and resolve rate is 50.9 points for Pine, 43.2 for Opus with Claude Code and 41.9 for GPT-5.6 Sol with Codex.
- [22]
Pine's published run data gives readers material to inspect, while Pine describes its broader 2-5x speed claim as based on preliminary internal tests that vary by task.
- [23]
Pine says its software helps an unnamed business automate audits and that the customer's team now handles 50% more work with the same people; the release does not explain how the increase was measured.
Sources
1 independent publisher whose own reporting we read for this story.
- runtimewire.comPine AI launches a cloud computer for agents working across business software
1 article · October 9, 2026
Topics and entities
Follow any of these and your For You feed starts watching them — no settings page required.