Best Text-to-speech APIs for Voice Agents (2026)
For voice agents, Cartesia is our pick (from $5/mo): For real-time voice agents, websocket streaming is critical for low-latency bidirectional conversation. Realtime speech for phone agents and voice bots, where latency and streaming decide whether the conversation feels human. Below is the full ranking and the tradeoffs, or read how we score.
If you sign up through links on this page, vsref may earn a commission; programs exist on both sides of most comparisons, and commissions never change verdicts. How we make money
What matters for voice agents
Weight ×5 = decisive, ×1 = relevant| Fact | Cartesia | Azure Speech | ElevenLabs | Fish Speech | Deepgram Aura-2 |
|---|---|---|---|---|---|
| TTFB latency (vendor-claimed)×5 | ~90 msJul 20 | n/a | ~75 msJul 20 | n/a | ~200 msJul 20 |
| Streaming audio output×5 | ✓ YesJul 20 | ✓ YesJul 20 | ✓ YesJul 20 | ✓ YesJul 20 | ✓ YesJul 20 |
| Realtime websocket API×4 | ✓ YesJul 20 | ✓ YesJul 20 | ✓ YesJul 20 | n/a | ✓ YesJul 20 |
| Price per 1M characters (flagship model)×3 | n/a | 22 $/1M charsJul 20 | 100 $/1M charsJul 20 | n/a | 30 $/1M charsJul 20 |
| Concurrency on base plan×3 | Free: 2Jul 20 | F0: 20 transactions per 60 secondsJul 20 | Free: 2 (Multilingual v2) / 4 (Flash)Jul 20 | n/a | up to 45 concurrent TTS requests (PAYG)Jul 20 |
The ranking, tool by tool
For real-time voice agents, websocket streaming is critical for low-latency bidirectional conversation. Cartesia supports a real-time websocket API while Mistral explicitly does not. Cartesia also claims 90 ms TTFB versus Mistral's 70 ms, but the websocket gap is decisive, since HTTP streaming alone cannot sustain the turn-taking required for phone agents. Cartesia additionally offers word-level timestamps and pronunciation dictionaries, both standard needs in production voice agent pipelines, while Mistral lacks both.
For voice agents, latency is decisive. Cartesia claims 90 ms TTFB versus Rime's 120 ms, a 25% advantage that directly shapes conversational feel. Both support real-time WebSocket APIs and streaming output, but Cartesia also offers official Python and JavaScript SDKs, making integration faster, while Rime has no official SDKs. Rime edges ahead on base concurrency (20 vs. 2 on the free tier) and HIPAA availability outside of enterprise plans, but those factors are secondary to raw latency for this use case. The latency gap tips the verdict to Cartesia.
For voice agents, latency is the decisive factor. Cartesia Sonic claims 90 ms TTFB, purpose-built for real-time conversation. Both tools offer websocket streaming and streaming audio output, but Cartesia also supports instant voice cloning, letting agents use custom brand voices, while OpenAI TTS has no cloning at all. Cartesia also covers 42 languages versus OpenAI's 13 fixed voices, giving broader deployment reach. The combination of ultra-low latency, custom voice capability, and a real-time websocket API makes Cartesia the clear winner for phone and voice bot use cases.
For voice agents, latency is the decisive factor. Cartesia (Sonic) claims 90 ms TTFB versus LMNT's 150 ms, a 40% advantage that directly determines conversational naturalness. Both support real-time WebSocket APIs and streaming output, so those features cancel out. LMNT offers no concurrency limits on paid plans, but Cartesia's speed edge matters more for real-time phone agents. Cartesia also supports 42 languages versus LMNT's 31, broadening deployment reach. The latency gap alone is enough to decide this clearly.
For voice agents, latency is decisive. Cartesia claims 90 ms TTFB versus Inworld's 200 ms, making Cartesia more than twice as fast to first audio byte, which directly determines how natural a conversation feels. Both tools offer real-time WebSocket APIs and streaming output, so those features are a wash. Cartesia also starts at $5 per month versus Inworld's $25 per month, lowering the barrier for prototyping voice bots. The latency gap alone is strong enough to tip the decision.
For voice agents, latency is decisive. Cartesia (Sonic) claims 90 ms TTFB versus ElevenLabs at 75 ms, so ElevenLabs is marginally faster. However, Cartesia offers self-hosting, which reduces real-world latency and gives infrastructure control critical for production voice bots. Cartesia also supports 42 languages versus ElevenLabs 32, broadening deployment reach. Both have realtime websocket APIs and streaming. The self-host advantage tips the verdict to Cartesia for voice agent deployments where consistent low latency and control matter most.
For voice agents, latency is the critical axis. Cartesia Sonic claims a TTFB of 90 ms versus Deepgram Aura-2's 200 ms, making Cartesia more than twice as fast to first audio byte. Both support real-time WebSocket APIs and streaming output. Cartesia also supports 42 languages versus Deepgram's 7, broadening deployment scenarios. Deepgram offers up to 45 concurrent PAYG requests, a meaningful operational advantage that keeps the margin from being clear-cut, but Cartesia's latency edge is the deciding factor for a conversation that feels human.
Both tools offer real-time WebSocket APIs and streaming output, so latency infrastructure is comparable. Azure pulls ahead on voice agent flexibility: it supports instant and professional voice cloning while OpenAI TTS supports neither. Azure also supports SSML, word-level timestamps, pronunciation dictionaries, and 100 languages, giving developers fine-grained control over agent speech. On cost, Azure flagship is 22 dollars per 1M chars versus OpenAI at 30 dollars, reducing per-call cost at scale. OpenAI TTS offers more output formats, but for voice agent deployments the cloning capability and lower cost tip the verdict to Azure.
For voice agents, real-time responsiveness is critical. Azure offers a WebSocket API enabling low-latency bidirectional streaming, while Google explicitly lacks one. Azure also supports 100 languages versus Google's 75, and its base plan allows 1000 concurrent paid connections versus Google's rate-limited default. The real-time WebSocket gap is the decisive differentiator for phone agent and voice bot deployments, where conversational feel depends on sub-second turn-taking.
For real-time voice agent use cases, WebSocket support is critical for low-latency bidirectional streaming. Azure offers a real-time WebSocket API while Amazon Polly explicitly does not. Azure also supports instant voice cloning, enabling consistent agent personas, and handles up to 64 KB of SSML per turn over WebSocket, versus Polly's 3,000-character request cap, which matters for longer agent turns. These two gaps make Azure clearly better suited for phone agent and voice bot deployments.
For real-time voice agents, the decisive feature is a WebSocket API for bidirectional streaming. ElevenLabs has a verified real-time WebSocket API while Mistral Voxtral TTS does not. Latency is nearly identical, with ElevenLabs at 75 ms versus Mistral at 70 ms, making that gap negligible. ElevenLabs also offers emotion and style controls plus pronunciation dictionaries for tuning agent personas, whereas Mistral lacks both. The missing WebSocket support in Mistral is a structural gap for phone-agent use cases requiring persistent low-latency duplex connections, giving ElevenLabs the edge despite Mistral's lower per-character price.
Both tools support real-time WebSocket API and streaming output, so baseline infrastructure is equal. ElevenLabs claims a 75 ms TTFB latency, a concrete low-latency figure directly relevant to conversational voice agents. On cost, ElevenLabs charges 50 dollars per 1M characters for its fast model versus OpenAI TTS at 15 dollars per 1M characters, making OpenAI cheaper. However, ElevenLabs counters with instant voice cloning and 3,000 voices compared to only 13 built-in voices for OpenAI, enabling persona-matched agents. The latency claim and voice flexibility tip the decision toward ElevenLabs for voice agent realism, though the cost gap keeps the margin narrow.
For real-time voice agents, latency is the decisive factor. ElevenLabs claims a TTFB of 75 ms versus Murf API's 130 ms, a nearly 2x advantage that directly affects conversational naturalness. Both tools support WebSocket streaming and real-time APIs, so infrastructure parity exists there. ElevenLabs also offers a larger voice library (3,000 voices vs. Murf API's 150), giving more persona flexibility for agent deployments. Murf API is cheaper per character (10 vs. 50 per 1M on fast models), which narrows the gap, but for voice agents the latency edge of ElevenLabs is the more critical differentiator.
For voice agents, a realtime WebSocket API is critical. ElevenLabs has a verified realtime WebSocket API while Google Cloud TTS does not. ElevenLabs also claims 75ms TTFB, enabling human-feeling latency. Both support streaming, but only ElevenLabs provides the low-latency WebSocket connection essential for interactive phone agents and voice bots. These two factors decisively favor ElevenLabs for this use case.
For real-time voice agents, latency is decisive. ElevenLabs claims a TTFB of 75 ms versus Fish Audio's 100 ms, a 25 ms advantage that meaningfully reduces conversational lag. Both tools offer WebSocket APIs and streaming output, so infrastructure parity exists there. ElevenLabs also adds word-level timestamps and pronunciation dictionaries, useful for agent tuning, plus SOC 2 Type II compliance that enterprise phone deployments often require. Fish Audio's base concurrency of 5 edges out ElevenLabs' free-tier concurrency of 2 to 4, but ElevenLabs' latency lead and richer agent-oriented feature set tip the verdict its way.
ElevenLabs has a vendor-claimed TTFB of 75 ms and a realtime WebSocket API, both critical for natural-feeling voice agent conversations. Dia/Dia2 is marked dormant and supports only 1 language versus ElevenLabs 32 languages, severely limiting deployment scope. ElevenLabs also provides Python and JavaScript SDKs and streaming output, forming a complete production-ready stack. Dia/Dia2 offers no realtime API equivalent in the facts, making it unsuitable for phone agent latency requirements.
Both tools offer real-time WebSocket APIs and streaming output, so core infrastructure is equal. ElevenLabs claims a 75 ms TTFB, which is a strong latency argument for conversational feel. It also provides 3,000 voices with emotion and style controls, giving richer persona options for voice agents. Azure wins on price at $22 versus $100 per 1M characters and supports 100 languages versus 32, which matters for multilingual deployments. However, for voice agents where human-like latency is the deciding factor, ElevenLabs' 75 ms TTFB tips the balance narrowly.
For voice agents, realtime responsiveness is decisive. ElevenLabs offers a realtime WebSocket API while Amazon Polly explicitly does not. ElevenLabs also claims a 75ms TTFB, making conversational turn-taking viable. Additionally, ElevenLabs supports 40,000 characters per request vs Polly's 3,000, giving more flexibility for longer agent utterances. Polly's lack of a WebSocket API is a fundamental architectural gap for phone agent use cases where continuous streaming connections are required.
Both tools support streaming output and instant voice cloning, which are baseline requirements for voice agents. Fish Speech streaming is verified while CosyVoice streaming is only vendor-claimed, giving Fish Speech a credibility edge on the most critical latency feature. Fish Speech also supports 80 languages versus CosyVoice's 9, which matters for global phone agent deployments. Fish Speech has a hosted API available, reducing deployment friction for production voice bots. CosyVoice's Apache-2.0 license is more permissive than Fish Speech's custom non-commercial license, which is a real commercial drawback, keeping this a narrow rather than clear decision.
For voice agents, latency is the decisive factor. Deepgram Aura-2 TTS claims 200 ms TTFB versus ElevenLabs at 75 ms, so ElevenLabs is faster on latency alone. However, Deepgram supports up to 45 concurrent TTS requests on PAYG while ElevenLabs free tier caps at 2 to 4 concurrent requests, which matters for production call volume. Deepgram also costs 30 dollars per 1M chars versus ElevenLabs at 100 dollars per 1M chars for flagship, a 3x cost advantage at scale. Both offer websocket streaming. For high-concurrency, cost-sensitive production voice agents, Deepgram's concurrency and pricing edge outweighs ElevenLabs' latency lead.
For voice agents, latency is decisive. Fish Audio claims a 100 ms TTFB versus MiniMax's 250 ms, a 2.5x advantage that directly determines conversational naturalness. Both support real-time WebSocket APIs and streaming output. Fish Audio also offers 83 languages versus 40 for MiniMax, which is useful for multilingual agent deployments. Its pure usage pricing without a platform fee is simpler at lower volumes. The margin is narrow because these TTFB figures are vendor-claimed and not independently verified, but the gap is large enough to favor Fish Audio.
For real-time voice agents, low latency and websocket streaming are critical. Murf claims 130 ms TTFB and offers a verified real-time websocket API purpose-built for conversational agents. Speechify supports streaming output but no websocket API is documented, meaning it likely relies on standard HTTP streaming only. Murf also allows 5 concurrent requests on the base plan and supports ALAW/ULAW output formats commonly used in telephony stacks, while Speechify outputs only MP3 and WAV. These factors give Murf a practical edge for phone agent deployments.
For voice agents, latency is decisive. Rime claims 120 ms TTFB versus Inworld TTS at 200 ms, a 40% advantage that directly impacts conversational feel. Rime also offers 20 concurrent streams on its base plan versus Inworld's 5, meaning it handles more simultaneous calls before throttling. Both support WebSocket streaming and real-time APIs, but the higher concurrency and lower latency tip the balance toward Rime for phone agent workloads.
For voice agents, low latency and real-time bidirectional streaming are critical. OpenAI TTS has a verified real-time WebSocket API, enabling true low-latency conversational turn-taking that phone agents require. Mistral Voxtral TTS lacks a real-time WebSocket API entirely. Mistral claims 70 ms TTFB, but that figure is vendor-claimed only and is irrelevant without a WebSocket loop. OpenAI also supports more output formats, including Opus, which is optimal for telephony. The WebSocket capability alone is decisive for voice bot use cases.
For realtime voice agents, websocket support is critical for low-latency bidirectional audio. OpenAI TTS offers a realtime websocket API while Google does not. Both support streaming output. On cost, Google's fast model is cheaper at 4 dollars per 1M chars versus OpenAI's fast model at 15 dollars per 1M chars, which slightly favors Google. However, realtime websocket capability is a hard architectural requirement for phone agents where turn-taking latency is decisive, giving OpenAI a clear functional edge that outweighs the cost difference.
For voice agents, two practical factors stand out. First, Google supports 1,000 requests per minute default concurrency versus Polly's 80 concurrent streams, meaning Google handles bursty call center traffic far better. Second, Google accepts 5,000 characters per request versus Polly's 3,000, reducing the need to split longer agent utterances mid-stream. Both tools lack a real-time websocket API and both support streaming output, so those factors cancel out. The concurrency gap is the decisive edge for production voice agent deployments.
Cloud-utility TTS at commodity prices. No won verdicts for this use case yet; it ranks on ties and near-misses.
Full-stack open TTS from Alibaba's speech team: 9 languages plus 18+ Chinese dialects, 150 ms streaming latency claim, instruction control of emotion/dialect/speed, and training + deployment scripts under Apache-2.0. No won verdicts for this use case yet; it ranks on ties and near-misses.
Best-known open model for scripted two-speaker dialogue rather than narration - Apache-2.0, English-only, GPU-oriented. No won verdicts for this use case yet; it ranks on ties and near-misses.
Cost-leader realtime TTS for voice agents and games; hybrid pay-as-you-go plus monthly credit plans that lower the per-1M-character rate as commitment grows; SOC 2 Type II with zero-data-retention posture. No won verdicts for this use case yet; it ranks on ties and near-misses.
Speed-and-affordability challenger: a single fast model (Blizzard 2) with streaming-first APIs, unlimited voice clones on every tier, and flat subscription pricing with per-1K overage. No won verdicts for this use case yet; it ranks on ties and near-misses.
Multilingual cloning-first TTS with aggressive pricing. No won verdicts for this use case yet; it ranks on ties and near-misses.
Developer platform spun out of the Speechify brand: transparent tiered pricing ($10-$499/mo plus per-1M overage), streaming-native Simba 3.2, and a bundled voice-agents product with flat per-minute rates. No won verdicts for this use case yet; it ranks on ties and near-misses.
Open-weights-friendly voice cloning TTS from a frontier AI lab. No won verdicts for this use case yet; it ranks on ties and near-misses.