Build1 distinct publisher3 min readPublished
Coasty Systems runs two agents through a user's task for free and keeps the paired trajectories and blind votes to license. The bet: published task sets decay faster than anyone can replace them.
The Engineer · Build desk

Compiled by The EngineerSomething wrong?How this is made
Submitting a task starts a short pipeline. Both agents get the same sandbox, the same action vocabulary and the same step budget, their positions are randomized, and vendor names are scrubbed before anyone votes [7]. The design leans on one rule more than any other. A battle whose visible output reveals which model produced it is thrown out of the rating calculation [8]. That is the correct way to defend a blind vote. It is also a filter that does not sample evenly, since the models it penalises are the ones with a habit of introducing themselves.
The thesis underneath came out of the founders' previous product. Coasty was a computer-use agent aimed at legacy desktop software and third-party portals [4], and it reported 82.81% on OSWorld Verified, a suite spanning browsers, office software and operating-system tools [5]. Jannu and Kovuru say building against that suite is what showed them the decay path: agents improve, developers optimize against the published list, and items work their way into training data [6].
The failure statistics deserve arithmetic. About a third of runs do not complete [1]. If the two agents in a battle failed independently of each other, both would fail in roughly 11.6% of battles [2]. The launch post reports about 7% [16], which sits 4.6 points below that baseline [3]. Task difficulty should push the observed figure above the independent product rather than below it, because a hard task tends to defeat both entrants. So either matchmaking is skewing which pairs meet, or completing a task and surviving a human judgement are being counted as different things. No battle counts and no per-model denominators were published [17], so a reader cannot settle which.
Whether any of these rates transfer depends on the incoming stream. The tasks are whatever users bring, such as researching flights or assembling an expense report [3], which means the test set is chosen by the people who wanted a free agent rather than by a model provider [9]. That is the attraction for a lab. It is also the caveat for anyone reading the ranking as a general capability measure: the sampling frame is one product's user base.
What is actually being sold is trace length. The public sample stops at voted battles and the first three actions of each run, while commercial tiers keep longer traces and additional battle metadata [11]. Three actions show that an agent opened a browser and typed something, but not where it got lost, the part a model developer would pay for. The corpus and the ranking are also governed separately: three earlier Coasty battles, one of them voted, sit in the underlying record while staying off the leaderboard [19]. Metering by trace length is defensible engineering. It also makes the free ranking the sales demo for the dataset.
Ranked by verification strength, evidence, and original report placement.
Both agents receive the same sandbox, action vocabulary and step budget; their positions are randomized and vendor names are scrubbed before judging.
A battle that exposes a model's identity in its visible output is excluded from the rating calculation.
CoArena did not attach battle counts or model-level denominators to those launch statistics, limiting comparisons with results from larger, fixed benchmarks.
Users receive the agents' output while Coasty Systems retains an expanding corpus for licensing, and each vote adds a preference label that could make the dataset more useful to model developers.
Prateek Jannu and Nitish Kovuru launched CoArena, a service that sends two computer-use agents through the same user-submitted task and asks people to judge the results without knowing which model produced them.
Users supply live work such as researching flights or assembling expense reports; CoArena runs the task twice, reveals the model identities after a vote and uses the preference data to update its leaderboard.
Distinct publishers with included, body-backed reporting in this cluster.
1 article · August 29, 2026
Follow any of these and your For You feed starts watching them — no settings page required.
build
App factory or agent fleet manager: the fork is whose rate limit stops the work1 distinct publisher
invest
Speed becomes a SKU: OpenAI and Google put a separate price on latency3 distinct publishers
build
PerceptionBench puts a number on the step your pipeline treats as free1 distinct publisher
build
Grok 4.6 lands in Copilot two days after launch, and the model picker becomes a procurement problem1 distinct publisher
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
One outlet, one launch post
Follow any figure in this story back and it ends at the same desk: the founders' launch post, their methodology post, or a page CoArena wrote about itself. Runtimewire, the only outlet here, says so plainly — it tags the revenue as company-supplied and notes the completion rates arrive with no battle counts. What is genuinely checkable is the design: the parity rules, the exclusion of identity-leaking battles, the licensing tiers, the policy keeping Coasty's own agent off the roster. What is not checkable is any performance or revenue number in the piece.
Two weeks old, currently paused
The service had been live about two weeks when the founders described a weekly doubling of an unnamed base and $60,000 of revenue that may or may not be CoArena's. Set against that: by August 30th the homepage says posting and judging are paused, no reason offered. The intake side — new tasks, new votes — is the part everything else in the business rests on, and it is the part that is off.
The thesis is ahead of the evidence
"Published task sets decay faster than anyone can replace them" is a good argument, and right now it is carrying more weight than anything underneath it. A ranking whose early votes came substantially from the people who submitted the tasks, reported without battle counts, with judging suspended a week after the launch post, is not yet the durable evaluation layer the framing suggests. One thing runs the other way: Runtimewire is scrupulous about the gaps, so the overstatement lives in the pitch rather than in the reporting of it.
The scorekeeper sells the scores
Coasty Systems ranks computer-use agents, licenses the preference data and trajectories that ranking produces, and builds a competing agent. It keeps that agent out of new matchmaking and off the ranked roster — the right call, and it still leaves the company grading the models of its likeliest customers, with three of its own earlier battles sitting in the underlying record. "The exhaust is the product" is the company's own phrase: free access exists because the tasks and votes are the inventory.
Clear mechanics, unverifiable scale
We can describe how this thing works with real precision, because the design is documented and internally consistent — parity conditions, blinding, voided leaks, a refitted Bradley-Terry rating, provisional status under 30 battles, tiered data access. We cannot verify a single quantity, identify one licensee, or say why judging stopped. High confidence about the structure and the conflict; low confidence about the size.