Skip to content

benchmark

Agents' Last Exam

Agent evaluation where GLM-5.3 scored 28.5, up from 23.8.

Known aliases

  • ALE

Relationships

No evidence-backed relationships are recorded.

Current stories

product8 publishers

OpenAI halves the API price of Sol and Luna against GPT-5.6's promotional rates

OpenAI says better caching and inference let it cut API prices for Sol and Luna by half, and the cost advantage it claims for the cheap tier over the old top tier comes in at one tenth on the benchmark it published and one hundredth in its summary.

Perspective Coverage

8 publishers
Builder
Builder 36%
Operator
Operator 42%
Investor
Investor 22%

Reality

Evidence40
Adoption
Insufficient
Hype gap+35
Incentives70
Confidence60
product1 publisher

OpenAI cuts prices on new GPT-6 Sol and Luna models

Sol now bills $2 and $10 per million tokens and Luna $0.10 and $0.50, while OpenAI quotes its own benchmark results per task, where the cheap model lands 2.2 points behind Sol on the software engineering test.

Publishers:thenextweb.com

Reality

Evidence42
Adoption55
Hype gap+25
Incentives78
Confidence52
product1 publisher

Astra cuts the computer-use task from about 75 minutes to 40

OpenAI has put its paused computer-use model into a few customers' hands, where the speed gain arrives alongside a reasoning trail outside investigators say is harder to follow. The containment work now sits with the customer.

Publishers:zdnet.com

Reality

Evidence34
Adoption22
Hype gap+38
Incentives70
Confidence38