Real-Time Voice AI Latency Benchmarks: Cartesia, Deepgram, and ElevenLabs Compared
Latency is the make-or-break metric for real-time voice AI applications. When Time-to-First-Byte (TTFB) exceeds 200ms, the natural flow of conversation falters, making interactions feel clunky and unnatural. To address this critical need, a recent benchmark study aggregated median latency and pricing metrics across leading streaming Text-to-Speech (TTS) APIs. The tests utilized WebSocket connections, focusing on US-East endpoints and averaging results across 1,000 requests per provider.
The findings highlight significant differences in performance and cost, crucial for developers building everything from sophisticated conversational agents to responsive voice assistants. The study specifically examined the performance of Cartesia's Sonic-3, Deepgram's Aura-2, and ElevenLabs' Flash v2.5 models.
Performance Metrics: TTFB Latency
The primary differentiator for real-time voice applications is the speed at which the first byte of synthesized speech is returned. This TTFB directly impacts the perceived responsiveness of the system.
- Cartesia (Sonic-3): Achieved a median TTFB of 85ms. This superior performance positions it as excellent for applications requiring the fastest possible turn-taking, minimizing delays between user input and AI response.
- Deepgram (Aura-2): Recorded a median TTFB of 115ms. While slower than Cartesia, this latency is still well within acceptable ranges for many real-time applications and represents a very good performance level.
- ElevenLabs (Flash v2.5): The benchmark shows a TTFB of 195ms. This latency, while still functional, approaches the critical 200ms threshold where natural conversational flow begins to degrade. It marks this model as adequate but less ideal for highly interactive, rapid-fire dialogue compared to the other two.
The surprise here is not just the performance difference, but the clear segmentation. Cartesia appears to have engineered its Sonic-3 model with an explicit focus on minimizing this initial latency, likely through architectural choices optimized for streaming and rapid response generation. Deepgram's Aura-2 offers a compelling balance, and ElevenLabs' Flash v2.5, while slower, may offer other advantages not captured in this specific latency test.
Pricing Structures: Cost Per Million Characters
Beyond raw speed, cost is a significant factor for large-scale deployments. The benchmark also evaluated pricing per one million characters, a standard unit for TTS services.
- Deepgram (Aura-2): Priced at $15.00 per 1 million characters. This makes it the most cost-effective option among the three for bulk processing.
- Cartesia (Sonic-3): Priced at $20.00 per 1 million characters. It is more expensive than Deepgram but offers a premium for its speed.
- ElevenLabs (Flash v2.5): Priced at $22.00 per 1 million characters. This positions it as the highest-priced option in this specific comparison, which, combined with its higher latency, raises questions about its value proposition for pure real-time performance needs.
Real-Time Suitability Assessment
Synthesizing the latency and pricing data, the study provides an overall assessment of each provider's suitability for real-time applications:
- Cartesia (Sonic-3): Rated Excellent. Its 85ms TTFB is the fastest, enabling the most natural and immediate conversational turn-taking.
- Deepgram (Aura-2): Rated Very Good. It offers a strong balance between solid performance (115ms TTFB) and the lowest bulk cost ($15/M chars), making it a highly competitive choice.
- ElevenLabs (Flash v2.5): Rated Adequate. While functional, its 195ms TTFB is close to the limit for seamless real-time interaction. Its higher price point further suggests it might be better suited for applications where latency is less critical, or where its unique voice cloning or stylistic features are paramount.
If you are building a voice application where every millisecond counts for user experience, Cartesia's Sonic-3 appears to be the current leader. However, Deepgram presents a compelling alternative that balances performance with cost efficiency. The question remains for developers: does ElevenLabs' Flash v2.5 offer unique features or voice quality that justify its higher latency and cost in specific niche applications?

Implications for Developers and Businesses
The choice of a real-time voice AI API has direct consequences on user experience, development costs, and overall application viability. For developers prioritizing immediate conversational feedback, Cartesia's 85ms TTFB is a significant advantage. This low latency is crucial for applications like interactive voice response (IVR) systems, real-time translation, and AI-powered customer support where delays can lead to user frustration and abandonment.
Deepgram's offering is particularly interesting for businesses focused on scaling their voice AI initiatives. Its combination of respectable latency and the lowest per-character cost makes it an attractive option for high-volume applications. This efficiency can translate into substantial cost savings as usage grows, without a crippling impact on user experience.
ElevenLabs, while slower in this specific real-time benchmark, is known for its high-quality synthetic voices and advanced voice cloning capabilities. This suggests that its Flash v2.5 model might be optimized for different use cases, perhaps where the expressiveness and naturalness of the voice itself are prioritized over absolute minimal latency. Developers might choose ElevenLabs for applications like audiobook narration, character voice generation in games, or personalized AI companions where voice fidelity is paramount, and a slightly longer response time is acceptable.
The benchmark underscores that there is no single
