AWS counted 13 SageMaker inference launches so far in 2026. The one that changes production behaviour most is a prioritized list of up to five instance types, each allowed its own model optimization settings.
Reality
- Evidence45
- Adoption
- Insufficient
- Hype gap+25
- Incentives85
- Confidence55
A single-box test on Ollama 0.34.0 shows a turn chained with previous_response_id coming back HTTP 200 and status completed at the same 41 input tokens as the same question sent with no history at all. The request struct has no field for the key.
Reality
- Evidence66
- Adoption26
- Hype gap−8
- Incentives22
- Confidence63
Hetzner's experimental endpoint serves one Qwen model from German and Finnish data centres at no charge, and its own documentation tells users to keep production environments off it. Future token prices have not been published.
Reality
- Evidence40
- Adoption18
- Hype gap+30
- Incentives60
- Confidence50
The eleven-check suite comes from a gateway maintainer who wrote it to be pointed at their own endpoint as well as everyone else's, and the checks that matter most are the ones where a divergence still answers with success.
Reality
- Evidence42
- Adoption
- Insufficient
- Hype gap+12
- Incentives58
- Confidence45
Cross-Region inference for GPT-5.6 on Amazon Bedrock means capacity ceilings are now fixed by changing a profile prefix, not by changing models. The tradeoff is where your data gets processed.
Reality
- Evidence58
- Adoption20
- Hype gap+12
- Incentives86
- Confidence57