Jahanzaib

Eleven v4 Turbo Is Not a Drop In Swap for Your Voice Agent

ElevenLabs launched Eleven v4 and v4 Turbo with 150ms time to first speech and 90 languages. For voice agents, Turbo also means a new WebSocket, a per plan session cap and a discount that ends October 12.

Jahanzaib Ahmed
8 min read
ElevenLabs logo on a dark tile sending a copper waveform, for Eleven v4 Turbo voice agents

Eleven v4 Turbo, the low latency voice model ElevenLabs released on September 28, will not run on the WebSocket most custom voice agents use today. The launch coverage skipped that detail and led with a median time to first speech of about 150 milliseconds, more than 90 languages and the audio tags that make the new Eleven v4 family sound performed rather than read.

All of those numbers hold up against the docs. But getting Turbo into a live call changes your latency math, your protocol code and how many calls you can run at once, and the launch price that makes testing cheap ends on October 12.

What shipped

Two models went live the same day in ElevenAgents, ElevenCreative and the API, including on free accounts, according to the ElevenLabs announcement. Eleven v4 is the expressive flagship for narration, dubbing and dialogue. Eleven v4 Turbo carries the same architecture tuned for speed, and ElevenLabs says it was optimized together with its own agent platform.

~150msmedian time to first speech for v4 Turbo, network removed (launch post)
90+languages, up from 70 in v3 (TechCrunch)
10 secof audio for an instant voice clone (launch post)

The launch post claims the top spot on the Artificial Analysis voice leaderboard and a roughly 75% win rate in blind tests against Cartesia, Inworld and Gemini TTS models, with ties counted as half. TechCrunch reports that inline tags such as [laughs] can now be stacked, and that the biggest quality jumps came in Japanese, Brazilian Portuguese, Mandarin and Cantonese.

ElevenLabs' launch video, mostly audio samples: the fastest way to judge the expressiveness claim yourself.

Price per 1,000 characters, before and after October 12

Both new models launched at 72% off until October 12, per the ElevenAPI pricing page, which also gives the v3 Conversational latency. The other latency figures are median inference from the models page and the launch post, and all of them exclude application and network time.

ModelList price per 1K charactersUntil Oct 12Median inferenceLanguagesRealtime socket
Flash v2.5$0.04$0.04~75ms32Text to Speech
v3 Conversational$0.04$0.04~280ms70+Text to Dialogue
Eleven v4 Turbo$0.04$0.011~100ms90+Text to Dialogue
Eleven v4$0.08$0.022not stated90+Text to Dialogue

ElevenLabs' pricing table treats 1,000 characters as roughly one minute of speech. Take a five minute support call where the agent talks for about half of it: around 2,500 characters. On Turbo at list price that is about $0.10 of speech per call, the same as Flash. On the launch price it is under three cents. Per 1,000 calls that is $100 against $27.50, and the discount lasts two weeks: enough to pay for a test.

My read: price is a wash from October 13 onward. The real trade is Flash's roughly 25 millisecond inference edge and 40,000 character request limit against v4's expressiveness and nearly three times the language coverage. For a bilingual receptionist in a language outside Flash's 32, that list may decide it on its own. For an English line that mostly reads back appointment times, it probably does not.

Where the 150 millisecond figure stops

The ~150ms is a median, measured over WebSocket streaming, with network latency "measured and removed for all systems", according to the launch post's footnote. The comparison set is other vendors (Cartesia Sonic 3.6, xAI, Gemini 3.8 Flash-Lite TTS and OpenAI's GPT-4o mini TTS), not ElevenLabs' own Flash. It is a fair number for the model. Your callers experience something longer.

ElevenLabs' own latency explainer spells out the gap: public internet round trips of 20 to 200 milliseconds, player buffers of around 500 milliseconds being common, and Professional Voice Clones running slower than default or instant clones. Your caller also waits on speech recognition and your LLM before a single character reaches the voice model. That full chain is your latency budget, and Turbo only owns one link of it.

Voice agent turn pipeline from caller speech through speech to text, LLM tokens, the dialogue buffer, Eleven v4 Turbo and network back to the caller
The benchmark times one card in this chain; the dialogue buffer in front of it waits for about eight words of LLM output unless you flush it.

There is one more link, and it is specific to v4. The realtime dialogue socket that serves it holds text until it has roughly 40 characters and 8 words before it emits audio, unless you send a flush. If your LLM streams tokens straight into the socket, the first audio cannot start until the model has written about a clause. I would flush at the first sentence or clause boundary rather than let the server wait, because that wait lands on the caller as silence.

Eleven v4 Turbo needs a different WebSocket

Switching takes more than a new model ID. According to ElevenLabs' comparison of its two realtime sockets, the Text to Speech WebSocket, the one most custom agent stacks use for Flash, takes no eleven_v3 or eleven_v4 models at all. Eleven v4 Turbo streams only over the Text to Dialogue WebSocket, which has a different URL, a different message shape and different rules.

  • New endpoint. wss://api.elevenlabs.io/v1/text-to-dialogue/stream-input with model_id=eleven_v4_turbo, instead of the per voice /v1/text-to-speech/{voice_id}/stream-input.
  • New messages. The first frame must register voices; after that you send inputs objects, each naming a voice_id and an optional new_turn.
  • One voice. eleven_v4_turbo accepts exactly one registered voice per connection. Plain eleven_v4 allows up to ten.
  • Fixed 20 second timeout. The socket closes after 20 seconds without a client message unless you send keep_alive. A caller reading out a long account history can easily talk for longer than that.
  • No multi context mode. ElevenLabs lists multi context orchestration, for barge in and parallel utterances, as a reason to stay on the Text to Speech socket. On the dialogue socket you need your own plan for cutting audio when a caller interrupts.

Here is the shape of the conversation, adapted from the realtime dialogue guide:

// connect
wss://api.elevenlabs.io/v1/text-to-dialogue/stream-input?model_id=eleven_v4_turbo&output_format=mp3_44100_128

// first frame: register the single voice (key via xi-api-key header or here)
{"voices": ["VOICE_ID"]}

// each chunk of LLM output; flush at a clause boundary to skip the buffer wait
{"inputs": [{"text": "Your appointment is on Friday at ten. ", "voice_id": "VOICE_ID", "new_turn": true}], "flush": true}

// during long caller turns, stop the 20 second timer
{"keep_alive": true}

If you build on ElevenAgents itself, ElevenLabs runs this connection for you. If you run your own STT, LLM and TTS pipeline, or a platform that calls ElevenLabs on your behalf, check which socket it opens before you promise anyone the new voice. My Retell AI vs Vapi breakdown shows how much of that plumbing each platform hides.

Every open dialogue socket counts against your plan

Capacity is metered differently on the dialogue socket, and this is where I would slow down. On the Text to Speech socket, only active generation counts against your plan's concurrency. On the dialogue socket, every open connection holds a "dialogue session" for its whole lifetime, whether audio is playing or not, per the concurrency section of the models page. Open one more than your plan allows and the connection is rejected with too_many_concurrent_requests.

PlanDialogue sessionsFlash concurrency
Starter216
Creator3510
Pro7020
Scale and Business10530

The session numbers look bigger, but they measure a different thing. ElevenLabs' rule of thumb is that a concurrency limit of 5 "can typically support up to approximately 100 simultaneous audio broadcasts" on the standard socket. My read: the ratio works because a call only generates audio part of the time. If it holds, a Pro plan's 20 Flash slots stretch to hundreds of calls. The same plan holding one dialogue socket per call stops at 70. If the evening peak on a round the clock receptionist line ever passes that, sort out a plan change or a connection strategy first.

Who should test before October 12

Test now if you serve languages Flash does not cover, if your agent handles emotional calls like billing disputes or medical intake, or if you sit on ElevenAgents and the socket work is not yours. At 72% off, a week of real traffic costs about a quarter of list price.

Wait if your stack is built on the Text to Speech socket with multi context barge in, if your peak concurrency is near your plan's session count, or if your latency budget already has no room. Flash is still faster on inference and, from October 13, costs the same.

  1. Measure your own time to first audio. Log it from your server on today's model so you have a baseline that includes network and buffering.
  2. Check which socket your stack opens. If it is /v1/text-to-speech/…/stream-input, budget for the dialogue socket changes above.
  3. Pull your peak concurrent calls for the last 30 days and compare them with your plan's dialogue sessions.
  4. Run a split test before October 12 on one line or one language, with flush at clause boundaries and keep_alive on long caller turns.

For the same benchmark versus real call gap on another stack, see my take on Google's Gemini 3.8 Live benchmarks. And if you would rather have the agent built and monitored for you, that is the work I do: see how I build inbound voice agents.

Frequently asked questions

How much does Eleven v4 Turbo cost?

On the ElevenAPI pricing page, Eleven v4 Turbo lists at $0.04 per 1,000 characters, discounted 72% to $0.011 until October 12, 2026. ElevenLabs treats 1,000 characters as roughly a minute of speech. After the promotion, Turbo costs the same per character as Flash v2.5 and v3 Conversational, while the full Eleven v4 model lists at $0.08.

Is Eleven v4 Turbo faster than Flash v2.5?

Not on ElevenLabs' own numbers. The models page lists Flash v2.5 at about 75 milliseconds of median inference and v4 Turbo at about 100 milliseconds, both excluding network and application time. Turbo's pitch is v4 expressiveness at close to Flash speed. In a live call the difference is usually smaller than network, speech recognition and LLM delays.

Can I use Eleven v4 Turbo on the Text to Speech WebSocket?

No. ElevenLabs' documentation says the Text to Speech WebSocket does not accept eleven_v3 or eleven_v4 models. Eleven v4 Turbo streams over the Text to Dialogue WebSocket, which uses a different endpoint and message format, allows one registered voice for Turbo, closes after 20 seconds of inactivity without keep_alive, and counts each open connection as a dialogue session.

Is Eleven v4 Turbo available in ElevenAgents?

Yes. ElevenLabs says both Eleven v4 and Eleven v4 Turbo are available in ElevenAgents, ElevenCreative and the API, including on free accounts. The ElevenAgents plan limits on concurrent calls still apply, for example 20 on the $99 Pro plan.

Feed to Claude or ChatGPT

Published

September 29, 2026

Category

Voice AI
Jahanzaib Ahmed

Jahanzaib Ahmed

AI Systems Engineer & Founder

AI Systems Engineer with 126 production systems shipped. I run AgenticMode AI (AI agents, RAG systems, voice AI) and ECOM PANDA (ecommerce agency). I build AI that works in the real world for businesses across home services, healthcare, ecommerce, SaaS, and real estate.