Best Text-to-speech APIs for Voice Agents (2026)
For voice agents, Cartesia is our pick (from $5/mo): For real-time voice agents, websocket streaming is critical for low-latency bidirectional conversation. Realtime speech for phone agents and voice bots, where latency and streaming decide whether the conversation feels human. Below is the full ranking and the tradeoffs, or read how we score.
If you sign up through links on this page, vsref may earn a commission; programs exist on both sides of most comparisons, and commissions never change verdicts. How we make money
Reviewed by vsref Editorialfacts verified Sep 23, 2026Methodology →
What matters for voice agents
Weight ×5 = decisive, ×1 = relevant| Fact | Cartesia | ElevenLabs | Azure Speech | Fish Speech | Murf API |
|---|---|---|---|---|---|
| TTFB latency (vendor-claimed)×5 | ~90 msJul 20 | ~75 msJul 20 | n/a | n/a | ~130 msJul 20 |
| Streaming audio output×5 | ✓ YesJul 20 | ✓ YesJul 20 | ✓ YesJul 20 | ✓ YesJul 20 | ✓ YesJul 20 |
| Realtime websocket API×4 | ✓ YesJul 20 | ✓ YesJul 20 | ✓ YesJul 20 | n/a | ✓ YesJul 20 |
| Price per 1M characters (flagship model)×3 | n/a | 100 $/1M charsJul 20 | 22 $/1M charsJul 20 | n/a | 30 $/1M charsJul 20 |
| Concurrency on base plan×3 | Free: 2Jul 20 | Free: 2 (Multilingual v2) / 4 (Flash)Jul 20 | F0: 20 transactions per 60 secondsJul 20 | n/a | 5Jul 20 |
The ranking, tool by tool
For real-time voice agents, websocket streaming is critical for low-latency bidirectional conversation. Full Cartesia vs Voxtral TTS verdict →
For voice agents, latency is decisive. Full Cartesia vs Rime verdict →
For voice agents, latency is the decisive factor. Full Cartesia vs OpenAI TTS verdict →
For voice agents, latency is the decisive factor. Full Cartesia vs LMNT verdict →
For voice agents, latency is decisive. Full Cartesia vs Inworld TTS verdict →
For voice agents, latency is decisive. Full Cartesia vs ElevenLabs verdict →
For voice agents, latency is the critical axis. Full Cartesia vs Deepgram Aura-2 verdict →
For real-time voice agents, the decisive feature is a WebSocket API for bidirectional streaming. Full ElevenLabs vs Voxtral TTS verdict →
Both tools support real-time WebSocket API and streaming output, so baseline infrastructure is equal. Full ElevenLabs vs OpenAI TTS verdict →
For real-time voice agents, latency is the decisive factor. Full ElevenLabs vs Murf API verdict →
For voice agents, a realtime WebSocket API is critical. Full ElevenLabs vs Google Cloud TTS verdict →
For real-time voice agents, latency is decisive. Full ElevenLabs vs Fish Audio verdict →
ElevenLabs has a vendor-claimed TTFB of 75 ms and a realtime WebSocket API, both critical for natural-feeling voice agent conversations. Full ElevenLabs vs Dia / Dia2 verdict →
Both tools offer real-time WebSocket APIs and streaming output, so core infrastructure is equal. Full ElevenLabs vs Azure Speech verdict →
For voice agents, realtime responsiveness is decisive. Full ElevenLabs vs Amazon Polly verdict →
Both tools offer real-time WebSocket APIs and streaming output, so latency infrastructure is comparable. Full Azure Speech vs OpenAI TTS verdict →
For voice agents, real-time responsiveness is critical. Full Azure Speech vs Google Cloud TTS verdict →
For real-time voice agent use cases, WebSocket support is critical for low-latency bidirectional streaming. Full Azure Speech vs Amazon Polly verdict →
Both tools support streaming output and instant voice cloning, which are baseline requirements for voice agents. Full Fish Speech vs CosyVoice verdict →
For voice agents, latency is the deciding factor. Full Murf API vs Azure Speech verdict →
For real-time voice agents, low latency and websocket streaming are critical. Full Murf API vs Speechify API verdict →
For voice agents, latency is the decisive factor. Full Deepgram Aura-2 vs ElevenLabs verdict →
For voice agents, latency is decisive. Full Fish Audio vs MiniMax Speech verdict →
For voice agents, latency is decisive. Full Rime vs Inworld TTS verdict →
For voice agents, low latency and real-time bidirectional streaming are critical. Full OpenAI TTS vs Voxtral TTS verdict →
For realtime voice agents, websocket support is critical for low-latency bidirectional audio. Full OpenAI TTS vs Google Cloud TTS verdict →
For voice agents, two practical factors stand out. Full Google Cloud TTS vs Amazon Polly verdict →
Cloud-utility TTS at commodity prices. No won verdicts for this use case yet; it ranks on ties and near-misses.
Full-stack open TTS from Alibaba's speech team: 9 languages plus 18+ Chinese dialects, 150 ms streaming latency claim, instruction control of emotion/dialect/speed, and training + deployment scripts under Apache-2.0. No won verdicts for this use case yet; it ranks on ties and near-misses.
Best-known open model for scripted two-speaker dialogue rather than narration - Apache-2.0, English-only, GPU-oriented. No won verdicts for this use case yet; it ranks on ties and near-misses.
Cost-leader realtime TTS for voice agents and games; hybrid pay-as-you-go plus monthly credit plans that lower the per-1M-character rate as commitment grows; SOC 2 Type II with zero-data-retention posture. No won verdicts for this use case yet; it ranks on ties and near-misses.
Speed-and-affordability challenger: a single fast model (Blizzard 2) with streaming-first APIs, unlimited voice clones on every tier, and flat subscription pricing with per-1K overage. No won verdicts for this use case yet; it ranks on ties and near-misses.
Multilingual cloning-first TTS with aggressive pricing. No won verdicts for this use case yet; it ranks on ties and near-misses.
Developer platform spun out of the Speechify brand: transparent tiered pricing ($10-$499/mo plus per-1M overage), streaming-native Simba 3.2, and a bundled voice-agents product with flat per-minute rates. No won verdicts for this use case yet; it ranks on ties and near-misses.
Open-weights-friendly voice cloning TTS from a frontier AI lab. No won verdicts for this use case yet; it ranks on ties and near-misses.