AWS puts prompt caching's saving at up to 90 percent on cache hits; its own ten-question example nets about 75 percent, and only while every request lands inside the five-minute default TTL.
Reality
- Evidence55
- Adoption
- Insufficient
- Hype gap+18
- Incentives80
- Confidence58
LiteLLM's own Terminal-Bench run puts a gpt-5.4-mini classifier routing across Haiku, Sonnet and Opus at $14.34 against Opus-5's $19.74 for the same 16 solved tasks. The saving works out at 26 cents a task.
Publishers:docs.litellm.ai
Reality
- Evidence45
- Adoption15
- Hype gap+35
- Incentives80
- Confidence48
The San Francisco startup records production work and expert corrections, turns them into task-specific tests, and reroutes traffic only when a cheaper candidate passes; the studies behind that are its own.
Reality
- Evidence38
- Adoption10
- Hype gap+30
- Incentives75
- Confidence45
Anthropic's prompt caching docs describe an automatic mode that keeps the breakpoint on the last cacheable block of every request. The five-minute clock starts when the request starts, so generation time counts against it.
Reality
- Evidence74
- Adoption20
- Hype gap+10
- Incentives68
- Confidence70
GitHub reports cost cuts of 67% on TerminalBench 2.1 and 36% on DeepSWE, all measured by GitHub, while the preview as described hands platform teams no way to set the routing policy or see which model wrote what.
Reality
- Evidence32
- Adoption15
- Hype gap+33
- Incentives74
- Confidence42
The Copilot team put a token-shortening utility through its agentic benchmarks, watched task cost rise even as individual responses shrank, and shipped a compressor that only touches install, build, test and progress output.
Reality
- Evidence44
- Adoption58
- Hype gap−14
- Incentives68
- Confidence52
AWS, NVIDIA and Heidi report cutting production speech-recognition inference cost by 75% by sharing one GPU through CUDA MPS instead of running a single model instance. How much of that saving transfers depends on the latency envelope you accept.
Reality
- Evidence28
- Adoption20
- Hype gap+38
- Incentives78
- Confidence36
The router picks a cheaper model for work that resembles work it has already seen, and falls back to the strongest one when nothing matches. The same traces that enable that call are the proposed training set.
Reality
- Evidence34
- Adoption17
- Hype gap+41
- Incentives74
- Confidence52
The two-week Enzyme removal is credible enough, but it is priced against a five-year staffing plan written by people who did not want the project, and Airbnb's published numbers let you check the rate that plan assumed.
Reality
- Evidence48
- Adoption58
- Hype gap+52
- Incentives76
- Confidence54
AWS, NVIDIA and Heidi Health report holding sub-second transcription while cutting 16 GPU instances to four. Per-GPU throughput rose only about 1.5x, so the rest of that saving came out of provisioning headroom.
Reality
- Evidence52
- Adoption44
- Hype gap+34
- Incentives82
- Confidence58
Auto now picks the model for each task in every Replit account, and getting a specific model back means a Core or Pro plan plus a mode that can charge usage credits. Stripe's reported $8 billion for OpenRouter prices the same layer.
Reality
- Evidence34
- Adoption46
- Hype gap+31
- Incentives78
- Confidence52
Scenematic's out-of-distribution gate sends its least understood prompts straight to the expensive render. The logic holds up; the harness meant to prove it still simulates the scores.
Reality
- Evidence52
- Adoption15
- Hype gap+18
- Incentives62
- Confidence45
A dev.to hands-on replaces the classifier Lambda with a Step Functions Choice state. The trade is a second Bedrock invocation on every question, including the cheap ones.
Reality
- Evidence58
- Adoption
- Insufficient
- Hype gap+28
- Incentives62
- Confidence52
A dev.to writeup documents why cache_read_input_tokens falls to zero mid-run: a cache_control breakpoint walks back at most 20 content blocks, and parallel tool calls clear that in one turn.
Reality
- Evidence28
- Adoption
- Insufficient
- Hype gap+27
- Incentives34
- Confidence33
Upstage is selling tool-calling discipline rather than reasoning, and says 370 billion tokens moved through OpenRouter in its first week. The price, the part that matters most, is still qualitative.
Reality
- Evidence27
- Adoption40
- Hype gap+33
- Incentives83
- Confidence41
IBM Research ran self-mined guidelines across eight models on AppWorld. One model gained 16.1 points for 5 percent more tokens; another gained nothing at all.
Reality
- Evidence52
- Adoption
- Insufficient
- Hype gap+15
- Incentives70
- Confidence45
A new beta fingerprints each request and names the first structural divergence from a prior response id. Until now the only signal was cache_read_input_tokens going to zero.
Reality
- Evidence62
- Adoption
- Insufficient
- Hype gap+14
- Incentives74
- Confidence55
A payments company is paying model-lab money for the layer that decides which model gets the request. The routing decision and the settlement decision are converging into one stack.
Reality
- Evidence24
- Adoption58
- Hype gap+42
- Incentives66
- Confidence29
GenRec beat a heavily tuned production recommender by 0.006 percent on 10 percent of traffic for four weeks. The online gain is trivial; the labeling economics are not.
Reality
- Evidence56
- Adoption24
- Hype gap+8
- Incentives62
- Confidence52