Skip to content

benchmark

OSWorld

OSWorld is a benchmark that evaluates AI agents on real computer tasks across executable desktop and OS environments, measuring end-to-end task completion.

Known aliases

  • OSWorld 2.0
  • OSWorld-2.0
  • OSWorld 2.0 offline
  • OSWorld 2.1
  • OSWorld-2.1

Relationships

No evidence-backed relationships are recorded.

Current stories

build11 publishersConfirmed

Prompt length sets what Anthropic's Haiku 5.5 price cut is worth to each workload

Anthropic launched Claude Haiku 5.5 at an average price about 75% below Haiku 4.5. The saving varies widely with prompt length, so teams moving classification, support or query traffic need to price their own requests before they switch models.

Perspective Coverage

8 publishers
Builder
Builder 53%
Operator
Operator 28%
Investor
Investor 19%

Reality

Evidence68
Adoption30
Hype gap+20
Incentives65
Confidence70
build1 publisherOne report

H Company's Holo4 merges its desktop and API specialists into one open-weight model

H Company's open-weight Holo4 27B handles GUIs, code, MCP servers and REST APIs in one model, scoring 61.7% on OSWorld 2.0 at an estimated $1.22 per task. The choice between specialists happens during training, so agent teams could drop the router from their stack if the scores hold on their own workloads.

Publishers:dev.to

Reality

Evidence35
Adoption
Insufficient
Hype gap+25
Incentives
Insufficient
Confidence40
product5 publishersConfirmed

OpenAI's 10-cent GPT-6.1 Sol price covers only cached input

OpenAI lists GPT-6.1 Sol at $2 per million input tokens and $10 per million output; the 10-cent figure in early coverage is its cached-input rate. Teams moving work off Astra should budget on the list rates and OpenAI's per-task costs.

Perspective Coverage

5 publishers
Builder
Builder 44%
Operator
Operator 34%
Investor
Investor 22%

Reality

Evidence50
Adoption
Insufficient
Hype gap+35
Incentives70
Confidence60
build8 publishersConfirmed

Gemini 3.8 Flash's introductory price doubles on December 31, 2026

Google's third Flash release in six weeks keeps the $0.75/$3.75 rate card. But the model also spends more tokens per task. Both numbers in your cost model are moving before the price even changes.

Perspective Coverage

8 publishers
Builder
Builder 53%
Operator
Operator 29%
Investor
Investor 18%

Reality

Evidence58
Adoption35
Hype gap+22
Incentives72
Confidence62
product8 publishersConfirmed

OpenAI halves the API price of Sol and Luna against GPT-5.6's promotional rates

OpenAI says better caching and inference let it cut API prices for Sol and Luna by half, and the cost advantage it claims for the cheap tier over the old top tier comes in at one tenth on the benchmark it published and one hundredth in its summary.

Perspective Coverage

8 publishers
Builder
Builder 36%
Operator
Operator 42%
Investor
Investor 22%

Reality

Evidence40
Adoption
Insufficient
Hype gap+35
Incentives70
Confidence60
product1 publisherOne report

OpenAI cuts prices on new GPT-6 Sol and Luna models

Sol now bills $2 and $10 per million tokens and Luna $0.10 and $0.50, while OpenAI quotes its own benchmark results per task, where the cheap model lands 2.2 points behind Sol on the software engineering test.

Publishers:thenextweb.com

Reality

Evidence42
Adoption55
Hype gap+25
Incentives78
Confidence52
build1 publisherOne report

Anthropic prices its newer Sonnet a third below Sonnet 4.5

Sonnet 4.5 still leads GPT-5 on the coding leaderboards, and GPT-5 lists about 46 percent below it on a 5:1 token mix. Anthropic's current Sonnet undercuts both of Sonnet 4.5's list prices, and that complicates a routing plan built on the older pair.

Publishers:dev.to

Reality

Evidence40
Adoption20
Hype gap+15
Incentives55
Confidence45