Build1 distinct publisher3 min readUpdated
A Canadian clinic-receptionist vendor split-tested four TTS engines on live patient calls. The engine its own team ranked first in blind listening had the worst completion rate.
The Engineer · Build desk
Compiled by The EngineerSomething wrong?How this is made
Autor, which runs a voice AI receptionist called Loquent for dental and healthcare clinics in Canada, split-tested four text-to-speech engines across 12,247 live patient calls over eight weeks [1][5]. The engine its own team ranked first in blind audio comparisons every time, OpenAI's tts-1-hd, came last on call completion [3][11].
That matters because it invalidates the standard procurement method, which Autor admits it used originally: generate a few clips, play them in a quiet office, pick the one that sounds best [4]. Loquent has been in production for over a year, handling thousands of calls a month for bookings, insurance questions and after-hours triage, and completion rate was sitting around 74 percent [5][6].
The test design is the part worth copying. Each engine took roughly equal volume, randomly assigned at call start, with prompts, the Anthropic Claude conversation backbone, the Twilio telephony layer and the clinic mix held constant [7]. The arms were ElevenLabs Turbo v2.5 as incumbent, OpenAI tts-1-hd, Deepgram Aura, and a fourth entrant Autor says it cannot name because of an NDA [8]. Five outcomes were tracked: completion, time-to-first-hang-up, human transfer requests, repeat-caller behaviour, and a one-question SMS survey at a subset of clinics [9]. Naturalness and voice quality were deliberately not scored in isolation [10].
Completion came out at 81.2 percent for Deepgram Aura, 78.4 for ElevenLabs, 76.1 for the unnamed engine and 71.8 for OpenAI [11]. That is a 9.4 point spread between best and worst [1], which is a third fewer abandoned calls at the top of the range than at the bottom [2]. Median time-to-first-byte ran 180ms for Aura, 320ms for ElevenLabs, 410ms for the unnamed engine and 480ms for OpenAI [12]. Ranked by latency, the four engines land in exactly the reverse order of their completion rates [3]. Autor reports that the correlation between response latency and hang-up rate was stronger than any voice quality metric it looked at [13].
The mechanism Autor proposes is that latency reads to a caller as hesitation, and hesitation reads as malfunction [14]. Survey comments included patients saying the system "seemed confused" when it was in fact waiting for audio to generate [15]. The company puts the damage in the first three to five seconds of the call [16]. Human transfer requests follow the same ordering: 12.8 percent for Aura against 19.7 percent for OpenAI [17], a 6.9 point gap [4].
Two cautions. This is one vendor's instrumentation of its own product, against specific model versions, on its own stack [8][1]. And the arms are roughly 3,062 calls each [5]; the sampling error on a proportion near 80 percent at that sample size is about 0.7 points [6], so the 9.4 point spread is real but the 2.8 points between Aura and ElevenLabs is a thinner result than it looks in a bullet list.
The operational conclusion is narrow and useful. If you are buying speech synthesis for a phone line, the acceptance criterion is a production outcome metric measured through your own stack, plus a latency ceiling in the contract, not a listening session [7][13]. Run the candidates concurrently with everything else frozen, as Autor did [7].
What to watch: Autor says it isolated first-turn drop-off and describes the spread as dramatic, but the per-engine figures are not in the material we have [18]. Also worth watching whether the repeat-caller and SMS survey results track completion or diverge from it [9], and whether the higher-latency vendors close the time-to-first-byte gap, which would collapse the whole ranking [12].
Follow any of these and your For You feed starts watching them — no settings page required.
Ranked by verification strength, evidence, and original report placement.
In blind audio comparisons, Autor's team ranked OpenAI tts-1-hd first every time; the author describes it as producing objectively beautiful speech.
Autor tested four TTS engines across 12,247 calls over 8 weeks, split-testing real patient calls at dental and healthcare clinics using its production voice AI receptionist Loquent.
Autor originally picked its TTS engine by generating a few sample clips, playing them for the team, and choosing the one that sounded best in a quiet office.
Loquent has run in production for over a year, handling thousands of automated calls per month for healthcare and dental clinics across Canada, including booking appointments, answering insurance questions and after-hours triage.
Before the test, Loquent's completion rate (share of calls where the patient finished the full interaction rather than hanging up or asking for a human) was hovering around 74%.
Each engine handled roughly equal volume, randomly assigned at call start, with all other variables constant: same prompts, same Anthropic Claude conversation backbone, same Twilio infrastructure, same clinics.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Large but unverifiable single-vendor field test
The design is unusually concrete for a vendor blog — 12,247 randomised live calls over 8 weeks, a held-constant stack, five pre-declared metrics and per-engine point estimates whose implied sampling error (~0.7pp per arm) is small next to the 9.4-point spread. Against that: one publisher, the interested party, no raw data, no per-arm counts, no significance testing, one arm unnamed under NDA, latency confounded with engine identity, and no independent replication. The numbers are usable as a hypothesis and a test template, not as settled measurement.
Real production use, confined to one vendor
There is concrete deployment behind the claims: a receptionist product running over a year at Canadian clinics, thousands of calls a month, a completed four-engine live test, a production switch of primary TTS to Deepgram Aura, and a hybrid routing plus latency-budgeting architecture in service. All of it is one company's own stack, self-reported, with no named clinics, third-party deployments or other teams shown to have adopted the completion-rate buying criterion.
Findings modestly oversold beyond their setting
The reported numbers are presented plainly and the author explicitly narrows the question to task completion, which keeps the gap small. It is positive rather than zero because the headline generalisation — buy voice engines on completion rate, the demo-best engine loses — rests on one vendor's single vertical, one dated OpenAI model version, an NDA'd fourth arm and a causal latency story asserted without published statistics. 'Objectively beautiful speech' from an informal internal listening panel and 'the correlation was stronger than any voice quality metric' both reach further than the disclosed evidence.
Vendor-authored comparison with undisclosed relationships
The piece is written by Autor about Autor's own product, publishes a favourable operational arc (74% baseline to 83.6% with its own hybrid architecture), and ranks four commercial suppliers by name while withholding one under NDA. No commercial relationships, credits, discounts or partnership terms with ElevenLabs, OpenAI or Deepgram are disclosed, and no vendor was given a chance to respond. That is a strong incentive to publish a rigorous-looking result that also functions as thought-leadership marketing for a clinic receptionist product.
Directionally credible, single-source
The internal consistency is good — latency order exactly inverts completion order, transfer requests and first-turn drop-off move the same way, and the vendor changed production behaviour on the result — so the direction (latency hurts task completion on phone calls) is plausible. Confidence stays below the midpoint because everything traces to one interested author, no statistics or data are released, one arm is anonymous, and nothing in the cluster is independently corroborated or contested.
product
Twilio's memory API removes the profile database from voice AI, and adds two model hops1 distinct publisher
build
A RAG demo becomes a product at the tenant boundary, not the retriever1 distinct publisher
product
Adobe ships Firefly's audio tools with a licensing claim attached, not just better output2 distinct publishers
build
Voicebot amnesia is a telephony bug: FreeSWITCH's ESL socket, not the model1 distinct publisher
Distinct publishers with included, body-backed reporting in this cluster.
dev.to
1 article · August 17, 2026