Skip to content

benchmark

Terminal-Bench

Terminal-Bench is a benchmark that evaluates how well AI agents complete real-world tasks via command-line interactions in a terminal environment.

Known aliases

  • Terminal Bench
  • terminal-bench-2
  • Terminal-Bench 2.0
  • Terminal Bench 2.1
  • Terminal-Bench 2.1
  • Terminal-Bench 3.0
  • TerminalBench 3.0
  • Terminal Bench 4
  • Terminal-Bench 4
  • Terminal Bench 4.0
  • Terminal-Bench 4.0
  • Terminal-Bench v2.1

Relationships

No evidence-backed relationships are recorded.

Current stories

invest5 publishers

Cached context shrinks the discount from Anthropic's half-price Sonnet 5.5

Anthropic released Sonnet 5.5 at $2 and $10 per million input and output tokens, half the Opus 5.5 rate. How much a buyer saves by moving work down a tier depends on tokens burned per task and on cache reads priced identically on both models.

Perspective Coverage

5 publishers
Builder
Builder 46%
Operator
Operator 38%
Investor
Investor 16%

Reality

Evidence55
Adoption35
Hype gap+20
Incentives70
Confidence60
build14 publishers

One model string moves Vercel AI Gateway traffic to Claude Sonnet 5.5

Vercel's AI Gateway now routes Claude Sonnet 5.5 through a single model ID, according to a dev.to review of the week's releases. The benchmark and cost figures come only from that third-party review, so a team's own tests decide when regulated workloads move.

Perspective Coverage

14 publishers
Builder
Builder 47%
Operator
Operator 30%
Investor
Investor 23%

Reality

Evidence58
Adoption48
Hype gap+35
Incentives62
Confidence60
build3 publishers

Kimi K3 on its cheapest host undercuts Fireworks' Ember-1 despite a 23% cut in reasoning tokens

Fireworks' Ember-1 used 23% fewer reasoning tokens than Kimi K3 in The New Stack's tests, yet Kimi on the cheapest host would cost $1.96 to Ember's $2.48. Ember beats Fireworks' own Kimi rate and loses at the cheapest, so buyers have to price the host before the model.

Perspective Coverage

3 publishers
Builder
Builder 52%
Operator
Operator 30%
Investor
Investor 18%

Reality

Evidence55
Adoption30
Hype gap+25
Incentives70
Confidence58
product11 publishers

Sonnet 5.5 comes within two points of Opus 5.5 at half the per-token price

Anthropic says Sonnet 5.5 nearly ties Opus 5.5, 1,844 to 1,846, on an everyday-work benchmark while running more than 30% faster than Sonnet 5. For teams paying double per token for Opus, Sonnet becomes the sensible default, with Opus kept for long, ambiguous jobs.

Perspective Coverage

11 publishers
Builder
Builder 36%
Operator
Operator 43%
Investor
Investor 21%

Reality

Evidence45
Adoption50
Hype gap+25
Incentives65
Confidence55
build4 publishers

DeepSeek open-sources the harness, then raises the price of the model

Harness v0.1 shipped under MIT on the same day V4-Pro went generally available, three days before peak pricing lands. The lock-in it targets is the runtime, not the weights.

Perspective Coverage

4 publishers
Builder
Builder 51%
Operator
Operator 31%
Investor
Investor 18%

Reality

Evidence55
Adoption
Insufficient
Hype gap+25
Incentives70
Confidence58
build5 publishers

GLM-5.3 keeps GLM-5.2's base model and claims 50% more on coding: plan for shorter eval cycles

Z.ai says every gain in GLM-5.3 came from post-training on an unchanged base. If that holds, refresh cadence for self-hosted weights is set by RL runs, not pretraining runs.

Perspective Coverage

5 publishers
Builder
Builder 58%
Operator
Operator 33%
Investor
Investor 9%

Reality

Evidence40
Adoption30
Hype gap+35
Incentives70
Confidence55
build8 publishers

Gemini 3.8 Flash's introductory price doubles on December 31, 2026

Google's third Flash release in six weeks keeps the $0.75/$3.75 rate card. But the model also spends more tokens per task. Both numbers in your cost model are moving before the price even changes.

Perspective Coverage

8 publishers
Builder
Builder 53%
Operator
Operator 29%
Investor
Investor 18%

Reality

Evidence58
Adoption35
Hype gap+22
Incentives72
Confidence62
build6 publishers

Spark 1.3's index jump lands on the three tests that carry half the score

Meta's fourth Muse Spark in five months gains four points on Artificial Analysis' Intelligence Index, almost entirely in agentic tests, while two scores fall and the tier Meta benchmarked against rivals is still a limited preview.

Perspective Coverage

6 publishers
Builder
Builder 52%
Operator
Operator 26%
Investor
Investor 22%

Reality

Evidence68
Adoption25
Hype gap+30
Incentives65
Confidence70
invest3 publishers

xAI holds Grok's $2 token price for a model 40 Elo points behind Fable 5.1

Grok 4.7 arrived on Monday after five walked-back timelines, with 40% more parameters and Grok 4.6's list price intact. Cursor's own cost chart still puts its price per task above GPT-6 Astra and Claude Sonnet 5.

Perspective Coverage

3 publishers
Builder
Builder 42%
Operator
Operator 27%
Investor
Investor 31%

Reality

Evidence45
Adoption35
Hype gap+15
Incentives70
Confidence55
product5 publishers

Price cuts minutes apart send agent routing back to the spreadsheet

Anthropic took 20% off Opus 5.5 and OpenAI halved its two new GPT-6 tiers the same day. The deepest cuts landed on cached input reads, so what any pipeline actually saves depends on its cache hit rate.

Perspective Coverage

5 publishers
Builder
Builder 38%
Operator
Operator 37%
Investor
Investor 25%

Reality

Evidence55
Adoption20
Hype gap+25
Incentives70
Confidence60
product7 publishers

Box measured Claude Opus 5.5 using a third of the tokens Opus 5 needed

Anthropic priced Opus 5.5 tokens 20 percent below Opus 5 and raised five-hour usage limits by the same amount. The larger saving in the launch is a token count from one customer's evaluation of its own content.

Perspective Coverage

7 publishers
Builder
Builder 41%
Operator
Operator 38%
Investor
Investor 21%

Reality

Evidence45
Adoption30
Hype gap+35
Incentives60
Confidence60

Earlier coverage

  1. Anthropic cuts Opus 5.5 prices 20% on tokens, 60% on cache reads, citing fewer tokens burned for 40% total savings

    Invest · September 23, 2026 · 1 publisher

  2. Grok 4.7 buys 5.9 points of CursorBench at Grok 4.6's token price

    Invest · September 22, 2026 · 1 publisher

  3. Fixing the deployment target splits the flash-tier coding leaderboard into three winners

    Build · September 21, 2026 · 1 publisher

  4. Strands Harness keeps five subsystems local and routes one call to Bedrock

    Build · September 21, 2026 · 1 publisher

  5. A Berkeley scanning agent scores 100% on five AI agent benchmarks without solving a task

    Science · September 20, 2026 · 1 publisher

  6. Anthropic prices its newer Sonnet a third below Sonnet 4.5

    Build · September 20, 2026 · 1 publisher

  7. Swapping the harness under a fixed model kept pass rate within 8 points on 50 Terminal-Bench Pro tasks

    Build · September 19, 2026 · 1 publisher

  8. Artificial Analysis retries a provider safety error ten times before scoring the attempt zero

    Build · September 19, 2026 · 1 publisher

  9. Rule-based elision before summarization led on efficiency among context-management strategies

    Leadership · September 19, 2026 · 1 publisher

  10. Databricks finds the harness swings coding-agent cost more than the model does

    Product · September 18, 2026 · 1 publisher

  11. Real-SWE licenses private production codebases to score coding agents on real business tasks

    Build · September 18, 2026 · 1 publisher

  12. Vals put Hy4 Preview first among open-weight models on code migration at $3.41 a test

    Build · September 17, 2026 · 1 publisher

  13. Fireworks' own DeepSWE numbers put four coding models inside the noise band

    Product · September 17, 2026 · 1 publisher

  14. LangChain's Deep Agents offloads oversized tool output to a filesystem at 20,000 tokens

    Build · September 16, 2026 · 1 publisher

  15. Arize's cheapest model per finished task reliably solves only a fifth of the benchmark

    Leadership · September 15, 2026 · 1 publisher

  16. Amodei's pacing plan would put outside evaluators inside Anthropic with the right to publish

    Leadership · September 15, 2026 · 1 publisher

  17. DeepSeek's smallest model beats its own 28-day-old flagship on seven of eight shared scores

    Invest · September 14, 2026 · 1 publisher

  18. DeepSeek's V4.1-Flash reads a million-token prompt on 8B active parameters

    Leadership · September 12, 2026 · 1 publisher

  19. GLM-5.3-Flash buys seven retries for the price of one Kimi K3 call

    Build · September 11, 2026 · 1 publisher

  20. DeepSeek reroutes V4-Pro API traffic to a smaller model on September 14

    Product · September 11, 2026 · 1 publisher

  21. ByteDance's self-evolved agent harnesses gain 3.11 held-out points inside a 4.75-point noise band

    Build · September 10, 2026 · 1 publisher

  22. Cognition put a cost penalty inside SWE-2's reinforcement-learning objective

    Build · September 10, 2026 · 1 publisher

  23. DeepSeek's V4 preview cuts million-token KV cache to a tenth of V3.2's

    Leadership · September 8, 2026 · 1 publisher

  24. Huang tucks 400,000 incoming GPUs into three-word "AGI has arrived" post

    Product · September 7, 2026 · 1 publisher

  25. Abliteration.ai rents a refusal-stripped GLM-5.3 for five dollars a million tokens

    Build · September 6, 2026 · 1 publisher

  26. Post-training alone took GLM-5.3 from 4.6 to 28.3 on Terminal-Bench 3.0

    Build · August 28, 2026 · 8 publishers

  27. GLM-5.3-Flash benchmarks its tenth-of-the-price claim against its own predecessor

    Leadership · September 5, 2026 · 1 publisher

  28. Four leaderboards, four denominators: what you buy when you standardize on a coding agent

    Product · August 25, 2026 · 1 publisher

  29. Harness-Bench makes the scaffold a measured variable, and model-only scores lose meaning

    Build · August 23, 2026 · 1 publisher

  30. GLM-5.3 changed nothing but the training environments. That is the whole test.

    Build · August 19, 2026 · 3 publishers

  31. Cost per successful task, not per token: a 2,400-run benchmark reorders the model shortlist

    Leadership · August 18, 2026 · 1 publisher

  32. Three frontier launches in a day, all pitched on price. Open weights set the ceiling.

    Build · August 14, 2026 · 4 publishers