Modulate raised $25 million to put models that analyze raw call audio, not transcripts, in front of more developers. Its pitch to teams running voice agents is a separate layer that spots cloned voices and grades how agents handle callers.
Perspective Coverage
3 publishers
- Builder
- Builder 40%
- Operator
- Operator 33%
- Investor
- Investor 27%
Reality
- Evidence58
- Adoption45
- Hype gap+20
- Incentives65
- Confidence60
Google added lip-syncing video avatars to its Gemini 3.8 Live voice agents, available now to Gemini Enterprise customers in 97 languages. Early reviews called the faces creepy and the launch came with no customer evidence, so support teams will have to measure trust in their own queues.
Perspective Coverage
7 publishers
- Builder
- Builder 33%
- Operator
- Operator 50%
- Investor
- Investor 17%
Reality
- Evidence58
- Adoption
- Insufficient
- Hype gap+25
- Incentives62
- Confidence60
Google's two new voice models process speech and reasoning at the same time and can fire tool calls without pausing the conversation. Published scores put the stronger one top of a speech-to-speech index and lowest on banking tasks.
Perspective Coverage
3 publishers
- Builder
- Builder 40%
- Operator
- Operator 40%
- Investor
- Investor 20%
Reality
- Evidence45
- Adoption
- Insufficient
- Hype gap+30
- Incentives70
- Confidence55
In a 30-call test suite for a home services intake agent, guardrail G4 told the model to stop collecting fields the moment it heard a gas smell and never told it when to resume. The dispatcher got the result.
Reality
- Evidence60
- Adoption8
- Hype gap−10
- Incentives35
- Confidence50
Oathra scores an agent phone call with separate code that reads the callee's turns, and its maintainer spent this week fixing the two natural English confirmation shapes that the checker itself got wrong.
Reality
- Evidence52
- Adoption12
- Hype gap+12
- Incentives55
- Confidence55
Microsoft treats speech as its own agent kind, with a managed realtime orchestrator holding the socket open. A dev.to walkthrough sets out why turn detection and tool calls had to be restructured around a few hundred milliseconds.
Reality
- Evidence38
- Adoption
- Insufficient
- Hype gap+15
- Incentives42
- Confidence40
Across three clinics and a bit over 1,000 real conversations, that overflow rule makes the agent's own call log the first count of ring-outs these practices have had. It is also the only source for the 20 to 30 recovered hours a month.
Reality
- Evidence28
- Adoption24
- Hype gap+12
- Incentives82
- Confidence34
Reception bundles the voice model, a phone number and a calendar write into a self-serve plan with 75 included minutes a month. That works out at about 39 cents a minute, the rate a home-built inbound agent has to beat.
Reality
- Evidence48
- Adoption
- Insufficient
- Hype gap+12
- Incentives72
- Confidence55
When the acoustic path changes mid-playback, the browser's echo canceller stops cancelling and the leftover sound reaches the voice detector as speech. The fix was 300 extra milliseconds, applied only during playback.
Reality
- Evidence58
- Adoption22
- Hype gap−6
- Incentives32
- Confidence48
A voice avatar kept its microphone button lit while the server logged 0.00 volume, and the fix needed three failed hypotheses before a log showed the widget enabling a track that was about to be discarded.
Reality
- Evidence55
- Adoption20
- Hype gap−5
- Incentives30
- Confidence55
Fangchang Ma says the lag and the dead-eyed listening in voice products come from the handoffs between four models. Nuance's evidence is a demo it calls unfinished and a research preview promised for later this year.
Reality
- Evidence28
- Adoption4
- Hype gap+45
- Incentives78
- Confidence52
A client's request for an expressive voice pushed a team off VAPI and onto Twilio, Deepgram and Cartesia. The metric they chased was time-to-first-audio: how fast the first syllable arrives. Owning the pipeline means owning barge-in.
Reality
- Evidence32
- Adoption18
- Hype gap+12
- Incentives45
- Confidence35
Measured per call rather than averaged, one voice agent's outbound cache hit rate came in at less than half of inbound on identical code and prompts, and the cause was two variable fields sitting inside the cached prefix.
Reality
- Evidence40
- Adoption20
- Hype gap+25
- Incentives30
- Confidence45
Partial recognition results and speech the user never heard end up in the prompt that drives the next turn. A dev.to tutorial's fix is a reducer with exactly two commit points, and the cost is an adapter you write yourself.
Reality
- Evidence60
- Adoption
- Insufficient
- Hype gap−12
- Incentives45
- Confidence55
AWS says three integration decisions carry Natera's phlebotomy booking agent. The reported numbers are good, and they cover a narrower job than "booked appointment".
Reality
- Evidence40
- Adoption34
- Hype gap+18
- Incentives84
- Confidence56
AWS puts a Claude-powered Connect agent on a real phone number and reaches the restaurant backend over MCP. The unforgiving part is the pause the caller hears while that happens.
Reality
- Evidence46
- Adoption
- Insufficient
- Hype gap+28
- Incentives88
- Confidence55
A dev.to breakdown makes the case that turn detection, endpointing and barge-in are the hard parts of a voice product, and a model swap does not touch any of them.
Reality
- Evidence32
- Adoption
- Insufficient
- Hype gap+12
- Incentives
- Insufficient
- Confidence38
A practitioner's account argues live transfer, structured message and flat refusal are separate code paths, fired by four different triggers. That is a twelve-cell table, not a toggle.
Reality
- Evidence24
- Adoption
- Insufficient
- Hype gap+14
- Incentives42
- Confidence38
A builder's account decomposes a production voice minute into speech-to-text, tokens, speech synthesis, telephony and media infrastructure, then puts failed calls in the numerator.
Reality
- Evidence34
- Adoption
- Insufficient
- Hype gap+12
- Incentives
- Insufficient
- Confidence38
A Canadian clinic-receptionist vendor split-tested four TTS engines on live patient calls. The engine its own team ranked first in blind listening had the worst completion rate.
Reality
- Evidence38
- Adoption33
- Hype gap+22
- Incentives72
- Confidence41