Frontier models score at most 0.17 on a synthetic support-agent test of trusting only officially labeled claims; a two-line rule scores 1.00. A careful model that acts on rumors once their label is stripped makes the case for enforcing the check in harness code.
Reality
- Evidence45
- Adoption
- Insufficient
- Hype gap+25
- Incentives35
- Confidence40
A practitioner post on dev.to says a third of callers press zero or hang up in week one, and argues the transfer path should be built before any intent is wired, measured by the turn at which calls end.
Reality
- Evidence30
- Adoption
- Insufficient
- Hype gap+10
- Incentives60
- Confidence38
A no-code support agent called an endpoint it invented, was refused twice, and still told the customer it could see her charges. Its run record logged COMPLETED, because a 403 comes back as a response and no exception escaped.
Reality
- Evidence58
- Adoption10
- Hype gap−6
- Incentives68
- Confidence45
Salesforce named seven agents across sales, support, help desk and supply chain, and most of the new features are generally available now. The module it says lets Hunter plan across a multi-week deal ships in Hunter alone.
Reality
- Evidence30
- Adoption12
- Hype gap+30
- Incentives72
- Confidence42
Front's VP of engineering scores each agent on four axes plus three rates, and warns that questions generated from your own documentation are answerable by construction, so they grade the easy half of production traffic.
Reality
- Evidence33
- Adoption24
- Hype gap+11
- Incentives70
- Confidence44
The arrangement is unannounced and covers only select accounts, yet it lands in a market where Intercom and Zendesk already bill per resolution and most buyers still say they would rather meter consumption.
Reality
- Evidence44
- Adoption55
- Hype gap+24
- Incentives62
- Confidence42
Institute banking data cited in Forbes puts 2025-vintage firms at roughly six months to AI in the workflow, against more than six years for the 2019 cohort. The review habits used to arrive with the years.
Reality
- Evidence28
- Adoption38
- Hype gap+34
- Incentives72
- Confidence42
A field guide on dev.to describes a logistics agent that cleared 94% of test cases and 11% of 4,000 real tickets a day. The model was fine. Nobody designed the system around it.
Reality
- Evidence30
- Adoption18
- Hype gap+12
- Incentives62
- Confidence40
A published production loop for customer service agents shows where the cost really sits: not the model, but the classify-execute-confirm middle where a write hits a payment processor.
Reality
- Evidence34
- Adoption
- Insufficient
- Hype gap+28
- Incentives72
- Confidence44