Skip to content

Topic

Agentic Coding and Long-Horizon Tasks

Multi-day, tool-using software and ML infrastructure work as the frontier evaluation target.

Current stories

build2 publishers

Cognition reports 4.8x token throughput on Nvidia Vera Rubin in its own coding-agent test

Cognition says its SWE-2 model produced up to 4.8 times more total token throughput on Nvidia's Vera Rubin NVL72 than on GB200, in a test on CoreWeave. The company-run result points to more agent capacity per system, but whether long coding tasks get cheaper depends on what the new systems cost per hour.

Reality

Evidence40
Adoption30
Hype gap+35
Incentives85
Confidence60
build5 publishers

GLM-5.3 keeps GLM-5.2's base model and claims 50% more on coding: plan for shorter eval cycles

Z.ai says every gain in GLM-5.3 came from post-training on an unchanged base. If that holds, refresh cadence for self-hosted weights is set by RL runs, not pretraining runs.

Perspective Coverage

5 publishers
Builder
Builder 58%
Operator
Operator 33%
Investor
Investor 9%

Reality

Evidence40
Adoption30
Hype gap+35
Incentives70
Confidence55
build8 publishers

Gemini 3.8 Flash's introductory price doubles on December 31, 2026

Google's third Flash release in six weeks keeps the $0.75/$3.75 rate card. But the model also spends more tokens per task. Both numbers in your cost model are moving before the price even changes.

Perspective Coverage

8 publishers
Builder
Builder 53%
Operator
Operator 29%
Investor
Investor 18%

Reality

Evidence58
Adoption35
Hype gap+22
Incentives72
Confidence62
build6 publishers

Spark 1.3's index jump lands on the three tests that carry half the score

Meta's fourth Muse Spark in five months gains four points on Artificial Analysis' Intelligence Index, almost entirely in agentic tests, while two scores fall and the tier Meta benchmarked against rivals is still a limited preview.

Perspective Coverage

6 publishers
Builder
Builder 52%
Operator
Operator 26%
Investor
Investor 22%

Reality

Evidence68
Adoption25
Hype gap+30
Incentives65
Confidence70
product5 publishers

Price cuts minutes apart send agent routing back to the spreadsheet

Anthropic took 20% off Opus 5.5 and OpenAI halved its two new GPT-6 tiers the same day. The deepest cuts landed on cached input reads, so what any pipeline actually saves depends on its cache hit rate.

Perspective Coverage

5 publishers
Builder
Builder 38%
Operator
Operator 37%
Investor
Investor 25%

Reality

Evidence55
Adoption20
Hype gap+25
Incentives70
Confidence60
product7 publishers

Box measured Claude Opus 5.5 using a third of the tokens Opus 5 needed

Anthropic priced Opus 5.5 tokens 20 percent below Opus 5 and raised five-hour usage limits by the same amount. The larger saving in the launch is a token count from one customer's evaluation of its own content.

Perspective Coverage

7 publishers
Builder
Builder 41%
Operator
Operator 38%
Investor
Investor 21%

Reality

Evidence45
Adoption30
Hype gap+35
Incentives60
Confidence60

Earlier coverage

  1. Anthropic prices its newer Sonnet a third below Sonnet 4.5

    Build · September 20, 2026 · 1 publisher

  2. A compliance agent reads its permission table and stops one transition short of signing

    Build · September 19, 2026 · 1 publisher

  3. Thompson credits the harness for the agent jump that retired his bubble call

    Build · September 19, 2026 · 1 publisher

  4. Lauren Tan's 2,000 pull requests a month rest on an app that fits in one process

    Build · September 19, 2026 · 1 publisher

  5. Vals put Hy4 Preview first among open-weight models on code migration at $3.41 a test

    Build · September 17, 2026 · 1 publisher

  6. OpenAI's improvement loop compiles five traced runs into a rerunnable Promptfoo gate

    Build · September 15, 2026 · 1 publisher

  7. A four-hour Claude Code session landed 42 unreviewed commits on main

    Build · September 14, 2026 · 1 publisher

  8. Anthropic takes one-cent warrants on 51 million shares of Truth Social's host

    Invest · September 14, 2026 · 1 publisher

  9. A zero border radius rule caused more arguments with Claude Code than the 1C sync did

    Build · September 13, 2026 · 1 publisher

  10. Parlotype's localization gate landed before roughly 450 keys left the markup

    Build · September 11, 2026 · 1 publisher

  11. Overnight laptop runs took over most of one Rust developer's Opus coding work

    Build · September 10, 2026 · 1 publisher

  12. Anthropic broke an agent ceiling by making "is this design good?" a gradable question

    Leadership · September 10, 2026 · 1 publisher

  13. Anthropic sold about 8 percent of itself for $30 billion

    Leadership · September 6, 2026 · 1 publisher

  14. Anthropic writes the agent handoff into the repository instead of the context window

    Leadership · September 6, 2026 · 1 publisher

  15. Fable 5.1 doubles science benchmark score, cuts bug-hunt task time by 3.6 seconds

    Build · September 5, 2026 · 1 publisher

  16. A $28 agent run swapped BCD for base-2^64 limbs and built its own oracle

    Build · September 5, 2026 · 1 publisher

  17. Anthropic's own playbook moves the software bottleneck into the review queue

    Product · September 3, 2026 · 1 publisher

  18. One Find and Replace task emptied a five-hour Codex window in 24 minutes

    Build · August 29, 2026 · 1 publisher

  19. A 107-page specification turned Claude Code into three native Task Managers

    Build · August 28, 2026 · 1 publisher

  20. Anthropic ships a price dial with its new model, and that is now the buying decision

    Leadership · August 26, 2026 · 1 publisher

  21. JetBrains asked 15,000 developers how much code agents write. The answers add up to 112 percent

    Build · August 26, 2026 · 1 publisher

  22. The $200 AI seat is a subsidy: ceiling users pull 40x to 70x what they pay

    Invest · August 23, 2026 · 1 publisher

  23. A CEO's $1,000 weekend, and the auto-renew setting that made it possible

    Invest · August 22, 2026 · 1 publisher

  24. Slack Code makes coding agents taggable teammates. Your merge-approval policy is now overdue

    Product · August 20, 2026 · 4 publishers

  25. GLM-5.3 changed nothing but the training environments. That is the whole test.

    Build · August 19, 2026 · 3 publishers