Eleven v4 Turbo Is Not a Drop In Swap for Your Voice Agent
ElevenLabs launched Eleven v4 and v4 Turbo with 150ms time to first speech and 90 languages. For voice agents, Turbo also means a new WebSocket, a per plan session cap and a discount that ends October 12.

Table of Contents
Eleven v4 Turbo, the low latency voice model ElevenLabs released on September 28, will not run on the WebSocket most custom voice agents use today. The launch coverage skipped that detail and led with a median time to first speech of about 150 milliseconds, more than 90 languages and the audio tags that make the new Eleven v4 family sound performed rather than read.
All of those numbers hold up against the docs. But getting Turbo into a live call changes your latency math, your protocol code and how many calls you can run at once, and the launch price that makes testing cheap ends on October 12.
What shipped
Two models went live the same day in ElevenAgents, ElevenCreative and the API, including on free accounts, according to the ElevenLabs announcement. Eleven v4 is the expressive flagship for narration, dubbing and dialogue. Eleven v4 Turbo carries the same architecture tuned for speed, and ElevenLabs says it was optimized together with its own agent platform.
The launch post claims the top spot on the Artificial Analysis voice leaderboard and a roughly 75% win rate in blind tests against Cartesia, Inworld and Gemini TTS models, with ties counted as half. TechCrunch reports that inline tags such as [laughs] can now be stacked, and that the biggest quality jumps came in Japanese, Brazilian Portuguese, Mandarin and Cantonese.
Price per 1,000 characters, before and after October 12
Both new models launched at 72% off until October 12, per the ElevenAPI pricing page, which also gives the v3 Conversational latency. The other latency figures are median inference from the models page and the launch post, and all of them exclude application and network time.
| Model | List price per 1K characters | Until Oct 12 | Median inference | Languages | Realtime socket |
|---|---|---|---|---|---|
| Flash v2.5 | $0.04 | $0.04 | ~75ms | 32 | Text to Speech |
| v3 Conversational | $0.04 | $0.04 | ~280ms | 70+ | Text to Dialogue |
| Eleven v4 Turbo | $0.04 | $0.011 | ~100ms | 90+ | Text to Dialogue |
| Eleven v4 | $0.08 | $0.022 | not stated | 90+ | Text to Dialogue |
ElevenLabs' pricing table treats 1,000 characters as roughly one minute of speech. Take a five minute support call where the agent talks for about half of it: around 2,500 characters. On Turbo at list price that is about $0.10 of speech per call, the same as Flash. On the launch price it is under three cents. Per 1,000 calls that is $100 against $27.50, and the discount lasts two weeks: enough to pay for a test.
My read: price is a wash from October 13 onward. The real trade is Flash's roughly 25 millisecond inference edge and 40,000 character request limit against v4's expressiveness and nearly three times the language coverage. For a bilingual receptionist in a language outside Flash's 32, that list may decide it on its own. For an English line that mostly reads back appointment times, it probably does not.
Where the 150 millisecond figure stops
The ~150ms is a median, measured over WebSocket streaming, with network latency "measured and removed for all systems", according to the launch post's footnote. The comparison set is other vendors (Cartesia Sonic 3.6, xAI, Gemini 3.8 Flash-Lite TTS and OpenAI's GPT-4o mini TTS), not ElevenLabs' own Flash. It is a fair number for the model. Your callers experience something longer.
ElevenLabs' own latency explainer spells out the gap: public internet round trips of 20 to 200 milliseconds, player buffers of around 500 milliseconds being common, and Professional Voice Clones running slower than default or instant clones. Your caller also waits on speech recognition and your LLM before a single character reaches the voice model. That full chain is your latency budget, and Turbo only owns one link of it.

There is one more link, and it is specific to v4. The realtime dialogue socket that serves it holds text until it has roughly 40 characters and 8 words before it emits audio, unless you send a flush. If your LLM streams tokens straight into the socket, the first audio cannot start until the model has written about a clause. I would flush at the first sentence or clause boundary rather than let the server wait, because that wait lands on the caller as silence.
Eleven v4 Turbo needs a different WebSocket
Switching takes more than a new model ID. According to ElevenLabs' comparison of its two realtime sockets, the Text to Speech WebSocket, the one most custom agent stacks use for Flash, takes no eleven_v3 or eleven_v4 models at all. Eleven v4 Turbo streams only over the Text to Dialogue WebSocket, which has a different URL, a different message shape and different rules.
- New endpoint.
wss://api.elevenlabs.io/v1/text-to-dialogue/stream-inputwithmodel_id=eleven_v4_turbo, instead of the per voice/v1/text-to-speech/{voice_id}/stream-input. - New messages. The first frame must register
voices; after that you sendinputsobjects, each naming avoice_idand an optionalnew_turn. - One voice.
eleven_v4_turboaccepts exactly one registered voice per connection. Plaineleven_v4allows up to ten. - Fixed 20 second timeout. The socket closes after 20 seconds without a client message unless you send
keep_alive. A caller reading out a long account history can easily talk for longer than that. - No multi context mode. ElevenLabs lists multi context orchestration, for barge in and parallel utterances, as a reason to stay on the Text to Speech socket. On the dialogue socket you need your own plan for cutting audio when a caller interrupts.
Here is the shape of the conversation, adapted from the realtime dialogue guide:
// connect
wss://api.elevenlabs.io/v1/text-to-dialogue/stream-input?model_id=eleven_v4_turbo&output_format=mp3_44100_128
// first frame: register the single voice (key via xi-api-key header or here)
{"voices": ["VOICE_ID"]}
// each chunk of LLM output; flush at a clause boundary to skip the buffer wait
{"inputs": [{"text": "Your appointment is on Friday at ten. ", "voice_id": "VOICE_ID", "new_turn": true}], "flush": true}
// during long caller turns, stop the 20 second timer
{"keep_alive": true}
If you build on ElevenAgents itself, ElevenLabs runs this connection for you. If you run your own STT, LLM and TTS pipeline, or a platform that calls ElevenLabs on your behalf, check which socket it opens before you promise anyone the new voice. My Retell AI vs Vapi breakdown shows how much of that plumbing each platform hides.
Every open dialogue socket counts against your plan
Capacity is metered differently on the dialogue socket, and this is where I would slow down. On the Text to Speech socket, only active generation counts against your plan's concurrency. On the dialogue socket, every open connection holds a "dialogue session" for its whole lifetime, whether audio is playing or not, per the concurrency section of the models page. Open one more than your plan allows and the connection is rejected with too_many_concurrent_requests.
| Plan | Dialogue sessions | Flash concurrency |
|---|---|---|
| Starter | 21 | 6 |
| Creator | 35 | 10 |
| Pro | 70 | 20 |
| Scale and Business | 105 | 30 |
The session numbers look bigger, but they measure a different thing. ElevenLabs' rule of thumb is that a concurrency limit of 5 "can typically support up to approximately 100 simultaneous audio broadcasts" on the standard socket. My read: the ratio works because a call only generates audio part of the time. If it holds, a Pro plan's 20 Flash slots stretch to hundreds of calls. The same plan holding one dialogue socket per call stops at 70. If the evening peak on a round the clock receptionist line ever passes that, sort out a plan change or a connection strategy first.
Who should test before October 12
Test now if you serve languages Flash does not cover, if your agent handles emotional calls like billing disputes or medical intake, or if you sit on ElevenAgents and the socket work is not yours. At 72% off, a week of real traffic costs about a quarter of list price.
Wait if your stack is built on the Text to Speech socket with multi context barge in, if your peak concurrency is near your plan's session count, or if your latency budget already has no room. Flash is still faster on inference and, from October 13, costs the same.
- Measure your own time to first audio. Log it from your server on today's model so you have a baseline that includes network and buffering.
- Check which socket your stack opens. If it is
/v1/text-to-speech/…/stream-input, budget for the dialogue socket changes above. - Pull your peak concurrent calls for the last 30 days and compare them with your plan's dialogue sessions.
- Run a split test before October 12 on one line or one language, with flush at clause boundaries and
keep_aliveon long caller turns.
For the same benchmark versus real call gap on another stack, see my take on Google's Gemini 3.8 Live benchmarks. And if you would rather have the agent built and monitored for you, that is the work I do: see how I build inbound voice agents.
Frequently asked questions
How much does Eleven v4 Turbo cost?
On the ElevenAPI pricing page, Eleven v4 Turbo lists at $0.04 per 1,000 characters, discounted 72% to $0.011 until October 12, 2026. ElevenLabs treats 1,000 characters as roughly a minute of speech. After the promotion, Turbo costs the same per character as Flash v2.5 and v3 Conversational, while the full Eleven v4 model lists at $0.08.
Is Eleven v4 Turbo faster than Flash v2.5?
Not on ElevenLabs' own numbers. The models page lists Flash v2.5 at about 75 milliseconds of median inference and v4 Turbo at about 100 milliseconds, both excluding network and application time. Turbo's pitch is v4 expressiveness at close to Flash speed. In a live call the difference is usually smaller than network, speech recognition and LLM delays.
Can I use Eleven v4 Turbo on the Text to Speech WebSocket?
No. ElevenLabs' documentation says the Text to Speech WebSocket does not accept eleven_v3 or eleven_v4 models. Eleven v4 Turbo streams over the Text to Dialogue WebSocket, which uses a different endpoint and message format, allows one registered voice for Turbo, closes after 20 seconds of inactivity without keep_alive, and counts each open connection as a dialogue session.
Is Eleven v4 Turbo available in ElevenAgents?
Yes. ElevenLabs says both Eleven v4 and Eleven v4 Turbo are available in ElevenAgents, ElevenCreative and the API, including on free accounts. The ElevenAgents plan limits on concurrent calls still apply, for example 20 on the $99 Pro plan.
Published
September 29, 2026
Category
Voice AI
Jahanzaib Ahmed
AI Systems Engineer & Founder
AI Systems Engineer with 126 production systems shipped. I run AgenticMode AI (AI agents, RAG systems, voice AI) and ECOM PANDA (ecommerce agency). I build AI that works in the real world for businesses across home services, healthcare, ecommerce, SaaS, and real estate.
Related articles

Google's Voice Model Won the Benchmark. The One It Tells You to Use Placed Fifth.
A breakdown of what Google shipped in Gemini 3.8 Live, why the headline benchmark belongs to the sibling model, and what actually breaks in your voice agent the day you change the model string.
NewsVoice AIVoice AI
Virtual Receptionist Companies: What They Cost in 2026, and the AI Option Nobody Quotes You
A straight look at what the big receptionist services actually charge in 2026, why the per minute math gets expensive fast, and the AI setup I use to get most businesses the same coverage for far less.
GuideVoice AIVirtual Receptionist
GPT-6.1 Sol Halves the Cache Rate and Leaves Every Other Price Alone
OpenAI's GPT-6.1 Sol keeps GPT-6 Sol's $2 input and $10 output prices and cuts only cached input. Here is what that saves on real agent loops, and why Chat Completions users cannot just swap the model name.
NewsImplementationOpenAI
