Build1 publisher3 min readPublished
The demo-best voice engine finished last: 12,247 calls argue for buying on completion rate
A Canadian clinic-receptionist vendor split-tested four TTS engines on live patient calls. The engine its own team ranked first in blind listening had the worst completion rate.
The Engineer · Build desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction
What happened
- Autor tested four TTS engines across 12,247 calls over 8 weeks, split-testing real patient calls at dental and healthcare clinics using its production voice AI receptionist Loquent.
- In blind audio comparisons, Autor's team ranked OpenAI tts-1-hd first every time; the author describes it as producing objectively beautiful speech.
- Autor originally picked its TTS engine by generating a few sample clips, playing them for the team, and choosing the one that sounded best in a quiet office.
- Loquent has run in production for over a year, handling thousands of automated calls per month for healthcare and dental clinics across Canada, including booking appointments, answering insurance questions and after-hours triage.
- Before the test, Loquent's completion rate (share of calls where the patient finished the full interaction rather than hanging up or asking for a human) was hovering around 74%.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
Autor, which runs a voice AI receptionist called Loquent for dental and healthcare clinics in Canada, split-tested four text-to-speech engines across 12,247 live patient calls over eight weeks [1][5]. The engine its own team ranked first in blind audio comparisons every time, OpenAI's tts-1-hd, came last on call completion [3][11].
That matters because it invalidates the standard procurement method, which Autor admits it used originally: generate a few clips, play them in a quiet office, pick the one that sounds best [4]. Loquent has been in production for over a year, handling thousands of calls a month for bookings, insurance questions and after-hours triage, and completion rate was sitting around 74 percent [5][6].
The test design is the part worth copying. Each engine took roughly equal volume, randomly assigned at call start, with prompts, the Anthropic Claude conversation backbone, the Twilio telephony layer and the clinic mix held constant [7]. The arms were ElevenLabs Turbo v2.5 as incumbent, OpenAI tts-1-hd, Deepgram Aura, and a fourth entrant Autor says it cannot name because of an NDA [8]. Five outcomes were tracked: completion, time-to-first-hang-up, human transfer requests, repeat-caller behaviour, and a one-question SMS survey at a subset of clinics [9]. Naturalness and voice quality were deliberately not scored in isolation [10].
Completion came out at 81.2 percent for Deepgram Aura, 78.4 for ElevenLabs, 76.1 for the unnamed engine and 71.8 for OpenAI [11]. That is a 9.4 point spread between best and worst [1], which is a third fewer abandoned calls at the top of the range than at the bottom [2]. Median time-to-first-byte ran 180ms for Aura, 320ms for ElevenLabs, 410ms for the unnamed engine and 480ms for OpenAI [12]. Ranked by latency, the four engines land in exactly the reverse order of their completion rates [3]. Autor reports that the correlation between response latency and hang-up rate was stronger than any voice quality metric it looked at [13].
The mechanism Autor proposes is that latency reads to a caller as hesitation, and hesitation reads as malfunction [14]. Survey comments included patients saying the system "seemed confused" when it was in fact waiting for audio to generate [15]. The company puts the damage in the first three to five seconds of the call [16]. Human transfer requests follow the same ordering: 12.8 percent for Aura against 19.7 percent for OpenAI [17], a 6.9 point gap [4].
Two cautions. This is one vendor's instrumentation of its own product, against specific model versions, on its own stack [8][1]. And the arms are roughly 3,062 calls each [5]; the sampling error on a proportion near 80 percent at that sample size is about 0.7 points [6], so the 9.4 point spread is real but the 2.8 points between Aura and ElevenLabs is a thinner result than it looks in a bullet list.
The operational conclusion is narrow and useful. If you are buying speech synthesis for a phone line, the acceptance criterion is a production outcome metric measured through your own stack, plus a latency ceiling in the contract, not a listening session [7][13]. Run the candidates concurrently with everything else frozen, as Autor did [7].
What to watch: Autor says it isolated first-turn drop-off and describes the spread as dramatic, but the per-engine figures are not in the material we have [18]. Also worth watching whether the repeat-caller and SMS survey results track completion or diverge from it [9], and whether the higher-latency vendors close the time-to-first-byte gap, which would collapse the whole ranking [12].