Skip to content

Topic

Model routing and orchestration

Routers that make model selection a runtime decision, and the argument that value accrues to that layer as models commoditise.

Current stories

product1 publisher

OpenAI's Dots moves always-on agents from owned hardware into OpenAI's cloud

OpenAI launched Dots at DevDay 2026, personal agents that run on cloud computers OpenAI maintains and keep context across ChatGPT, Slack and Teams. For unattended agent work, teams now choose between OpenAI's machines and hardware they can switch off themselves.

Publishers:turingpost.com

Reality

Evidence35
Adoption
Insufficient
Hype gap+25
Incentives55
Confidence40
product1 publisher

Fireworks' own DeepSWE numbers put four coding models inside the noise band

The vendor selling the cheapest model in the comparison reports a 0.7-point quality spread across four frontier models against run-to-run variation of 1.4 to 3.2 points. That leaves price per task, $0.43 against an implied $6.45 for GPT-6 Astra.

Publishers:fireworks.ai

Reality

Evidence42
Adoption18
Hype gap+28
Incentives88
Confidence58

Earlier coverage

  1. A three-agent CrewAI run spent 44.6 of its 118 seconds inside coworker tool calls

    Build · September 14, 2026 · 1 publisher

  2. OpenRouter's US endpoint rejects any request it cannot decrypt and serve in-country

    Build · September 14, 2026 · 1 publisher

  3. Gemini 3.6 Flash bills output tokens at five times the input rate

    Build · September 13, 2026 · 1 publisher

  4. OpenRouter's default Fusion slug lets the calling model decide when to spend 5x

    Build · September 13, 2026 · 1 publisher

  5. Mixing self-hosted Qwen with Bedrock Claude costs you telemetry, not a rewrite

    Build · August 14, 2026 · 1 publisher

  6. Routing easy jobs to cheaper models got Uber more than nine times the AI usage

    Product · September 13, 2026 · 1 publisher

  7. OpenAI's new max image tier costs about 35 times its cheapest one

    Invest · September 13, 2026 · 1 publisher

  8. Runway's ARR doubled to $200M in five months, driven by enterprise growth

    Invest · September 8, 2026 · 1 publisher

  9. Uber halved the cost of an AI session by routing work away from frontier models

    Leadership · September 11, 2026 · 1 publisher

  10. Spotify's shunt hook blocks non-targeted Claude Code file reads past a configurable line threshold

    Build · September 8, 2026 · 1 publisher

  11. Anthropic and GitHub have moved AI costs from the seat to the meter

    Leadership · September 3, 2026 · 1 publisher

  12. Thirty-nine retries fit inside the price gap between GLM-5.3-Flash and Opus 4.8

    Build · August 31, 2026 · 1 publisher

  13. Microsoft retired four models from the Foundry router under every deployment left on defaults

    Build · August 31, 2026 · 1 publisher

  14. Judging the cheap model's output beats guessing which prompt is hard

    Build · August 30, 2026 · 1 publisher

  15. Coding agents cost $4,125 a month because 73% of it is context you already sent

    Build · August 23, 2026 · 1 publisher

  16. Fable 5 at $50 per million output tokens turns model routing into a budget line

    Build · August 23, 2026 · 2 publishers

  17. Anthropic's discounts leave a 7.5x-to-37.5x gap, and the routing code decides the rest

    Invest · August 23, 2026 · 1 publisher

  18. Cost per successful task, not per token: a 2,400-run benchmark reorders the model shortlist

    Leadership · August 18, 2026 · 1 publisher

  19. Your inference bill is an architecture defect: declare the task before you call the model

    Build · August 18, 2026 · 1 publisher

  20. Three frontier launches in a day, all pitched on price. Open weights set the ceiling.

    Build · August 14, 2026 · 4 publishers