Build3 publishers3 min readPublished Updated
ElevenLabs aims v4 Turbo at live voice agents with a self-measured 150 ms to first speech
ElevenLabs released Eleven v4 Turbo for voice agents, reporting medians of about 100 ms inference latency and 150 ms to first speech. Moving an existing agent takes more than a model ID swap, since older voice clones need retraining and SSML break tags no longer work.
The Engineer · Build desk

What happened
- Turbo adds bidirectional streaming, so developers can send text as a language model produces it and receive audio before the sentence is finished.
- Artificial Analysis' Provider Voice Arena ranks the standard Eleven v4 first, with an Elo score of 1,319 from 1,674 samples.
- ElevenLabs' X thread says an Instant Voice Clone can capture a voice from 10 seconds of audio, while its v4 documentation says a clone generally uses one to two minutes.
- A two-week API launch offer prices Eleven v4 at $22 and Turbo at $11 per million characters.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- decision Choosing a model for agent traffic means setting the arena-ranked standard v4 against Turbo's lower latency and half-price launch rate, and only a listening test on a team's own voices settles it.
- constraint A multilingual agent that must keep one speaker's accent across languages starts from the wrong default, because v4 steers cross-language cloned speech toward a fluent target-language accent.
- contradiction Runtimewire says v4 comes at no extra cost on eligible Creator+ plans within credit limits, while TestingCatalog says both models are on every plan including the free tier, so the cost of a small pilot depends on which report holds.
Turbo comes with two latency figures, and they measure different points in a request. The roughly 100 ms figure is ElevenLabs' reported median inference latency [2]. The 150 ms figure is the median time to first speech [3]. ElevenLabs is positioning Turbo for support, sales and scheduling agents [22]. On a call, time to first speech is closer to the pause a caller hears. It is also the figure ElevenLabs used against competitors [4]. Both are medians. Neither report includes tail percentiles, and an agent's slowest turns sit in the tail.
Reaching that figure inside an agent loop depends on the wiring. ElevenLabs offers streaming and non-streaming endpoints [23]. Bidirectional streaming on Turbo [5] lets synthesis start while the language model is still writing. An agent that buffers the full LLM reply before calling synthesis puts the reply's generation time in front of Turbo's 150 ms.
ElevenLabs' comparison puts Turbo 112 ms ahead of Cartesia Sonic 3.6 and 664 ms ahead of OpenAI's GPT-4o mini TTS [1]. ElevenLabs ran that test [4]. For the 112 ms lead to hold on production calls, a team's network path to each provider, its text chunk sizes and its voices would need to resemble what ElevenLabs measured. A 664 ms gap leaves room for setup differences. The 112 ms gap is the one I would re-measure on my own traffic before switching.
The quality evidence belongs to the other model. The arena ranking is for the standard Eleven v4 [6], and the arena scores each provider's native voices [7]. A cloned voice on Turbo, delivered as telephony-ready mu-law audio [21], is outside what it measured. Runtimewire reports that ElevenLabs' post cites the ranking without spelling out the methodology [8].
Swapping the model ID is a one-line change [15]. Retraining the older clones [12] is where ElevenLabs' two cloning figures collide. The documentation's usual Instant Voice Clone sample is six to twelve times the length the launch thread promised [3]. Both numbers come from ElevenLabs, about the same feature. I'd retrain against the documentation and treat the 10 seconds as a best case [9]. Professional Voice Clones, unavailable in v3, are back in v4 [13], and Turbo keeps the same Professional Voice Clone across a call [20].
Prompt code needs its own change. v4 disables SSML break tags and takes pacing from natural-language audio tags, the bracketed directions such as laughs and whispers that go straight into the script [14]. An agent that inserts `<break>` tags to space out a confirmation number has to move that pacing into audio tags.
At launch rates, a maximum-length generation of 10,000 characters [16] costs $0.22 on v4 and $0.11 on Turbo [5]. The offer runs for two weeks [17].
What to watch
- An independent measurement of Turbo's time to first speech that reports p95 and p99 latency alongside the median.
- The per-character rates for Eleven v4 and Turbo once the two-week API launch offer ends.
- Whether ElevenLabs revises the 10-second cloning claim or the one-to-two-minute sample guidance in its v4 documentation.