Skip to content

benchmark

Humanity's Last Exam

A multi-domain benchmark of about 2,500 graduate-level questions across subjects like chemistry, economics and literature, testing AI reasoning limits.

Known aliases

  • HLE
  • HLE w/ Tools
  • Humanity's Last Exam

Relationships

No evidence-backed relationships are recorded.

Current stories

build4 publishers

DeepSeek open-sources the harness, then raises the price of the model

Harness v0.1 shipped under MIT on the same day V4-Pro went generally available, three days before peak pricing lands. The lock-in it targets is the runtime, not the weights.

Perspective Coverage

4 publishers
Builder
Builder 51%
Operator
Operator 31%
Investor
Investor 18%

Reality

Evidence55
Adoption
Insufficient
Hype gap+25
Incentives70
Confidence58
product1 publisher

Fireworks' own DeepSWE numbers put four coding models inside the noise band

The vendor selling the cheapest model in the comparison reports a 0.7-point quality spread across four frontier models against run-to-run variation of 1.4 to 3.2 points. That leaves price per task, $0.43 against an implied $6.45 for GPT-6 Astra.

Publishers:fireworks.ai

Reality

Evidence42
Adoption18
Hype gap+28
Incentives88
Confidence58
leadership3 publishers

ARC Prize puts Astra 37 points below the score OpenAI led with

The same model produced 99.9% in OpenAI's launch post and 62.7% on the benchmark authors' neutral harness, and Astra's input tokens cost double GPT-5.6 Sol's, which leaves the vendor table doing very little work in a purchase decision.

Perspective Coverage

3 publishers
Builder
Builder 27%
Operator
Operator 37%
Investor
Investor 36%

Reality

Evidence66
Adoption32
Hype gap+61
Incentives79
Confidence71
build3 publishers

OpenAI puts latency on the price list: 750 tokens/sec, gated by workload fit

The Ultrafast preview runs GPT-5.6 Sol on Cerebras hardware for a hand-picked customer list. That makes capacity allocation, not model choice, the constraint your architecture has to survive.

Perspective Coverage

3 publishers
Builder
Builder 42%
Operator
Operator 33%
Investor
Investor 25%

Reality

Evidence42
Adoption24
Hype gap+32
Incentives78
Confidence58