Latency optimization | ElevenLabs Documentation

This guide covers the core principles for improving text-to-speech latency.

While there are many individual techniques, we’ll group them into four principles.

Four principles

Enterprise customers benefit from increased concurrency limits and priority access to our rendering queue. Contact sales to learn more about our enterprise plans.

Use Flash models

Flash models deliver ~75ms inference speeds, making them ideal for real-time applications. The trade-off is a slight reduction in audio quality compared to Multilingual v2.

75ms refers to model inference time only. Actual end-to-end latency will vary with factors such as your location & endpoint type used.

Leverage streaming

There are three types of text-to-speech endpoints available in our API Reference:

Regular endpoint: Returns a complete audio file in a single response.
Streaming endpoint: Returns audio chunks progressively using Server-sent events.
Websockets endpoint: Enables bidirectional streaming for real-time audio generation.

Streaming

Streaming endpoints progressively return audio as it is being generated in real-time, reducing the time-to-first-byte. This endpoint is recommended for cases where the input text is available up-front.

Streaming is supported for the Text to Speech API, Voice Changer API & Audio Isolation API.

Websockets

The text-to-speech websocket endpoint supports bidirectional streaming making it perfect for applications with real-time text input (e.g. LLM outputs).

Setting auto_mode to true automatically handles generation triggers, removing the need to manually manage chunk strategies.

If auto_mode is disabled, the model will wait for enough text to match the chunk schedule before starting to generate audio.

For instance, if you set a chunk schedule of 125 characters but only 50 arrive, the model stalls until additional characters come in—potentially increasing latency.

For implementation details, see the text-to-speech websocket guide.

Consider geographic proximity

We serve our models from multiple regions to optimize latency based on your geographic location. By default all self-serve users use our US region.

For example, using Flash models with Websockets, you can expect the following TTFB latencies via our US region:

Region	TTFB
US	150-200ms
EU	230ms*
North East Asia	250-350ms
South Asia	380-440ms

*European customers can access our dedicated European tech stack for optimal latency of 150-200ms. Contact your sales representative to get onboarded to our European infrastructure.

We are actively working on deploying our models in Asia. These deployments will bring speeds closer to those experienced by US and EU customers.

Choose appropriate voices

We have observed that in some cases, voice selection can impact latency. Here’s the order from fastest to slowest:

Default voices (formerly premade), Synthetic voices, and Instant Voice Clones (IVC)
Professional Voice Clones (PVC)

Higher audio quality output formats can increase latency. Be sure to balance your latency requirements with audio fidelity needs.

We are actively working on optimizing PVC latency for Flash v2.5.