Skip to content

BuildNot yet confirmed elsewhere1 publisher3 min readPublished

Pine AI's agent computer finishes fewer benchmark tasks than two rivals, by Pine's own count

Pine AI's cloud computer for agents resolved 27.4% of SaaS-Bench tasks in Pine's own test, below two rivals at 31.1% and 29.2%. Its higher step score and much lower token bill come from whole systems running different models, so the computer's own contribution has not been measured.

The Engineer · Build desk

How we use AISend a correction

Illustration accompanying Pine AI's agent computer finishes fewer benchmark tasks than two rivals, by Pine's own count
Generated illustration

What happened

  • The service is in private beta for developers building products that need AI to act across websites, files and business software.
  • On UniPat AI's SaaS-Bench v1.1, Pine reports a 78.3% checkpoint score for GPT-5.6 Luna on its computer. Opus 5 running in Claude Code scored 74.3%, and GPT-5.6 Sol running in Codex scored 71.1%.
  • Model-token cost came to about $1.02 per task for Pine's system, against $26.50 and $20.50 for the two alternatives, according to Pine.

Why it matters

  • contradiction Pine's system ranks first of three on steps passed and last on tasks finished, and a buyer paid on completed work will be judged by the second ranking.
  • constraint No run holds the model and budget fixed while swapping the computer, so the benchmark cannot yet test Wang's claim that the environment is what holds agents back.
  • cost Pine's token spend stays near $3.72 per resolved task even at its lower resolve rate, but the infrastructure Pine actually sells is left out of every cost figure.
  • exposure Products built on the SDK put their users' sign-ins and business systems inside Pine's sandbox and takeover flow, so Pine's isolation layer becomes part of each developer's security boundary.

Dylan Wang's premise comes from networking. In Pine's launch post dated October 5, the co-founder recalled his time at Agora, where he came to see that audio and video quality could be improved by rebuilding the network underneath [3]. Pine Computer applies that idea to agents. It changes the environment the model works in, so the model does not have to make up for a computer designed around human eyes and hands [4]. Wang's bet is that a more capable model will not be enough on its own, and that agents also need a computer built differently [2].

A conventional computer-use agent runs a loop. It looks at a screenshot, decides what to click, acts, and looks at another screenshot [5]. Pine says its environment returns structure instead: the elements on the page, the actions available, and what changed after the last action [6]. The screen is still there. A user can watch it live, take over for a sign-in or an approval, and hand control back [7]. I'd expect a page-state diff to cost fewer tokens per step than an image. I'd also expect it to turn the question of whether a click worked into a lookup instead of a visual guess. Pine's evidence cannot confirm either point yet.

SaaS-Bench v1.1, from UniPat AI, covers 106 workflows across 23 applications [10]. Its checkpoint score counts intermediate steps passed, not tasks completed end to end [11]. Pine leads the two rivals on checkpoints by 4.0 and 7.2 points [16]. It trails them on resolved tasks by 3.7 and 1.8 points [17]. The gap between its two scores is the widest of the three systems: 50.9 points, against 43.2 and 41.9 [21]. If the resolve rate is counted over the 106 workflows, the systems finished about 29, 33 and 31 tasks [18]. With gaps of four tasks and two, I would want repeated runs before ranking anyone.

Each row in the table changes more than one variable. Pine's entry runs GPT-5.6 Luna on Pine Computer. The rivals are Opus 5 in Claude Code and GPT-5.6 Sol in Codex [10]. Runtimewire notes that the comparisons cover complete systems with different software and budgets, and do not isolate the effect of Pine's computer [14]. The closest pair shares a version number. To test Wang's thesis, a run would have to hold the model and the budget fixed and swap only the computer.

Pine reports about $1.02 in model tokens per task, against $26.50 and $20.50 for the alternatives [13]. The rivals spent about 26 and 20 times as much [19]. Pine's lower resolve rate does not close that gap. Its tokens come to about $3.72 per resolved task, while the cheaper rival's spend stays above $65 per resolved task at either rival's resolve rate [20]. The figures exclude infrastructure [14]. Infrastructure is what Pine sells. It runs the computer, browser and isolation layer, and it packages them with the agent runtime and permission handling as one cloud service [8][15].

Pine published its run data, so the scores can be checked [22]. It also labels its broader 2-5x speed claim as preliminary internal testing that varies by task [22]. I'd count both in Pine's favor.

Developers adopt it through an SDK and build the user experience around Pine's machine [8]. Each computer runs in its own sandbox, and developers keep their own keys [9]. Tasks can run in parallel, including on sites that have no API [9]. Pine has given one deployment example, an unnamed customer that uses its software to automate audits. Pine says that customer's team now handles 50% more work with the same people. The release does not explain how the increase was measured [23].

What to watch

  • A SaaS-Bench run that holds GPT-5.6 Luna and the budget fixed while swapping Pine Computer for a screenshot harness.
  • Pricing for the computer layer as the private beta opens, to set beside the $1.02 token figure.
  • An independent rerun of Pine's published run data, with repeated runs on the roughly 30 tasks each system resolved.

Clarity's read

What the record supports and how the coverage leans. The claims behind it follow.

Reality

Evidence35
Adoption10
Hype gap+40
Incentives80
Confidence40
Why these scores

Claim ledger

Ranked by verification strength, evidence, and original report placement.

  1. [1]

    Pine AI introduced Pine Computer for developers building products that need AI to work across websites, files and business software; the product is in private beta.

    ReportedSupportedSource: runtimewire.com, citing Pine's PR Newswire releaseView cited source
  2. [2]

    Pine AI co-founder Dylan Wang is betting that agents need a different kind of computer, not just a more capable model.

    ReportedSupportedSource: runtimewire.comView cited source
  3. [3]

    In Pine's launch post, dated October 5th, Wang wrote that at Agora he saw how rebuilding the underlying network could improve audio and video quality.

    ReportedSupportedSource: runtimewire.com, describing Pine's launch postView cited source

Sources

1 independent publisher whose own reporting we read for this story.

  1. runtimewire.com

    1 article · October 9, 2026

    Pine AI launches a cloud computer for agents working across business software

Share your take

Let Clarity write the post for you.

Signed-in readers get a short post drafted on this story in the register they choose — narrative, analytical, or a direct position — editable to the last word before it goes anywhere. The share buttons at the top of this story work without an account.

Topics and entities

Follow any of these and your For You feed starts watching them — no settings page required.

Loading related stories