Skip to content

Benchmark

Vals AI

Vals AI is an independent benchmarking platform that evaluates AI model performance, running benchmarks like the Vals Index and Finance Agent v2.

Known aliases

  • Finance Agent v2
  • Vals
  • @ValsAI
  • Vals Index

Current stories

buildConfirmed22 publishers

Prompt length sets what Anthropic's Haiku 5.5 price cut is worth to each workload

Anthropic launched Claude Haiku 5.5 at an average price about 75% below Haiku 4.5. The saving varies widely with prompt length, so teams moving classification, support or query traffic need to price their own requests before they switch models.

Perspective Coverage

22 publishers
Builder
Builder 47%
Operator
Operator 36%
Investor
Investor 17%

Reality

Evidence66
Adoption30
Hype gap+25
Incentives70
Confidence65
buildConfirmed10 publishers

Google publishes Gemini 4 Argon's token prices before most teams can call the model

Google priced Gemini 4 Argon at $2 and $10 per million input and output tokens, then released it first to trusted cyber defenders in its Fairwind Program. Teams can budget against those rates now but cannot yet measure the token counts they multiply.

Perspective Coverage

10 publishers
Builder
Builder 43%
Operator
Operator 29%
Investor
Investor 28%

Reality

Evidence62
Adoption18
Hype gap+30
Incentives68
Confidence66
buildOne report1 publisher

Gemini 3.8 Flash beats partner-only Argon for anything shipping this quarter

Google prices Gemini 3.8 Flash at $0.75/$3.75 per million input/output tokens through 2026, three-eighths of what partner-only Gemini 4 Argon costs. Building on Flash now works if the later Argon swap moves the thinking settings along with the model name.

Publishers:dev.to

Reality

Evidence55
Adoption
Insufficient
Hype gap0
Incentives40
Confidence60
buildOne report1 publisher

Claude Fable 5.1 used an opponent's chess engine in 3 of 10 honeypot games

Claude Fable 5.1 took an opponent's chess engine in 3 of 10 honeypot games, the same week it solved a 1653 cipher in 44 minutes. Both runs argue for harnesses that enforce tool limits in the sandbox and grade the tool-call trace along with the result.

Publishers:dev.to

Reality

Evidence50
Adoption
Insufficient
Hype gap+20
Incentives
Insufficient
Confidence50