Skip to content

benchmark

AutomationBench v1.0.6

Agentic automation benchmark where GLM-5.3 is reported at 48.2 versus 26.2.

Known aliases

  • AutomationBench
  • AutomationBench 1.0.6
  • AutomationBench-AA
  • AutomationBench v1.0.6
  • Zapier AutomationBench

Relationships

No evidence-backed relationships are recorded.

Current stories

product5 publishers

OpenAI's 10-cent GPT-6.1 Sol price covers only cached input

OpenAI lists GPT-6.1 Sol at $2 per million input tokens and $10 per million output; the 10-cent figure in early coverage is its cached-input rate. Teams moving work off Astra should budget on the list rates and OpenAI's per-task costs.

Perspective Coverage

5 publishers
Builder
Builder 44%
Operator
Operator 34%
Investor
Investor 22%

Reality

Evidence50
Adoption
Insufficient
Hype gap+35
Incentives70
Confidence60
product4 publishers

Meta hires MongoDB's CEO to sell its weeks-old Muse agent to businesses

Meta hired CJ Desai, MongoDB's CEO of about 11 months, to run Meta Enterprise Platform, a new unit selling its Muse agent and AI stack to businesses. So far a buyer can weigh Desai's enterprise record and a product list, with no prices or contract terms.

Perspective Coverage

4 publishers
Builder
Builder 25%
Operator
Operator 36%
Investor
Investor 39%

Reality

Evidence62
Adoption
Insufficient
Hype gap+35
Incentives65
Confidence60
product5 publishers

Price cuts minutes apart send agent routing back to the spreadsheet

Anthropic took 20% off Opus 5.5 and OpenAI halved its two new GPT-6 tiers the same day. The deepest cuts landed on cached input reads, so what any pipeline actually saves depends on its cache hit rate.

Perspective Coverage

5 publishers
Builder
Builder 38%
Operator
Operator 37%
Investor
Investor 25%

Reality

Evidence55
Adoption20
Hype gap+25
Incentives70
Confidence60
product8 publishers

OpenAI halves the API price of Sol and Luna against GPT-5.6's promotional rates

OpenAI says better caching and inference let it cut API prices for Sol and Luna by half, and the cost advantage it claims for the cheap tier over the old top tier comes in at one tenth on the benchmark it published and one hundredth in its summary.

Perspective Coverage

8 publishers
Builder
Builder 36%
Operator
Operator 42%
Investor
Investor 22%

Reality

Evidence40
Adoption
Insufficient
Hype gap+35
Incentives70
Confidence60
invest16 publishers

Sol's 27-cent benchmark task undercuts Opus 5 by more than eleven times

Anthropic and OpenAI shipped cheaper model tiers minutes apart on Tuesday. OpenAI halved Sol's posted token prices, and the only per-task cost comparison between the two labs so far comes from OpenAI itself.

Perspective Coverage

16 publishers
Builder
Builder 29%
Operator
Operator 35%
Investor
Investor 36%

Reality

Evidence62
Adoption25
Hype gap+25
Incentives70
Confidence60
product1 publisher

OpenAI cuts prices on new GPT-6 Sol and Luna models

Sol now bills $2 and $10 per million tokens and Luna $0.10 and $0.50, while OpenAI quotes its own benchmark results per task, where the cheap model lands 2.2 points behind Sol on the software engineering test.

Publishers:thenextweb.com

Reality

Evidence42
Adoption55
Hype gap+25
Incentives78
Confidence52