vsref

Best Text-to-speech APIs for Voice Agents (2026)

For voice agents, Cartesia is our pick (from $5/mo): For real-time voice agents, websocket streaming is critical for low-latency bidirectional conversation. Realtime speech for phone agents and voice bots, where latency and streaming decide whether the conversation feels human. Below is the full ranking and the tradeoffs, or read how we score.

If you sign up through links on this page, vsref may earn a commission; programs exist on both sides of most comparisons, and commissions never change verdicts. How we make money

Reviewed by vsref Editorialfacts verified Sep 23, 2026Methodology →

1Cartesia logoCartesiaWINNER
Lowest-latency TTS for real-time voice agents10 of 14 points · 7 matchups
Premium AI voice platform for creators and developers11 of 20 points · 10 matchups
Enterprise hyperscaler TTS with custom-voice depth5 of 10 points · 5 matchups

What matters for voice agents

Weighted attribute comparison for Voice Agents
FactCartesiaElevenLabsAzure SpeechFish SpeechMurf API
TTFB latency (vendor-claimed)×5~90 msJul 20~75 msJul 20n/an/a~130 msJul 20
Streaming audio output×5✓ YesJul 20✓ YesJul 20✓ YesJul 20✓ YesJul 20✓ YesJul 20
Realtime websocket API×4✓ YesJul 20✓ YesJul 20✓ YesJul 20n/a✓ YesJul 20
Price per 1M characters (flagship model)×3n/a100 $/1M charsJul 2022 $/1M charsJul 20n/a30 $/1M charsJul 20
Concurrency on base plan×3Free: 2Jul 20Free: 2 (Multilingual v2) / 4 (Flash)Jul 20F0: 20 transactions per 60 secondsJul 20n/a5Jul 20
Swipe → to see every tool column.
×5 TTFB latency (vendor-claimed): Time to first audio byte is the single biggest driver of how responsive a voice agent feels.×5 Streaming audio output: Agents must start speaking before the full reply is synthesized; non-streaming APIs are a hard stop.×4 Realtime websocket API: A persistent websocket avoids per-request connection overhead in live conversations.×3 Price per 1M characters (flagship model): Per-character rate compounds fast at call-center volumes.×3 Concurrency on base plan: Concurrent call capacity on the entry plan decides when you are forced into a bigger contract.

The ranking, tool by tool

1Cartesia logoCartesiaWINNER
For real-time voice agents, websocket streaming is critical for low-latency bidirectional conversation.

For real-time voice agents, websocket streaming is critical for low-latency bidirectional conversation. Full Cartesia vs Voxtral TTS verdict →

For voice agents, latency is decisive. Full Cartesia vs Rime verdict →

For voice agents, latency is the decisive factor. Full Cartesia vs OpenAI TTS verdict →

For voice agents, latency is the decisive factor. Full Cartesia vs LMNT verdict →

For voice agents, latency is decisive. Full Cartesia vs Inworld TTS verdict →

For voice agents, latency is decisive. Full Cartesia vs ElevenLabs verdict →

For voice agents, latency is the critical axis. Full Cartesia vs Deepgram Aura-2 verdict →

For real-time voice agents, the decisive feature is a WebSocket API for bidirectional streaming.

For real-time voice agents, the decisive feature is a WebSocket API for bidirectional streaming. Full ElevenLabs vs Voxtral TTS verdict →

Both tools support real-time WebSocket API and streaming output, so baseline infrastructure is equal. Full ElevenLabs vs OpenAI TTS verdict →

For real-time voice agents, latency is the decisive factor. Full ElevenLabs vs Murf API verdict →

For voice agents, a realtime WebSocket API is critical. Full ElevenLabs vs Google Cloud TTS verdict →

For real-time voice agents, latency is decisive. Full ElevenLabs vs Fish Audio verdict →

ElevenLabs has a vendor-claimed TTFB of 75 ms and a realtime WebSocket API, both critical for natural-feeling voice agent conversations. Full ElevenLabs vs Dia / Dia2 verdict →

Both tools offer real-time WebSocket APIs and streaming output, so core infrastructure is equal. Full ElevenLabs vs Azure Speech verdict →

For voice agents, realtime responsiveness is decisive. Full ElevenLabs vs Amazon Polly verdict →

Both tools offer real-time WebSocket APIs and streaming output, so latency infrastructure is comparable.

Both tools offer real-time WebSocket APIs and streaming output, so latency infrastructure is comparable. Full Azure Speech vs OpenAI TTS verdict →

For voice agents, real-time responsiveness is critical. Full Azure Speech vs Google Cloud TTS verdict →

For real-time voice agent use cases, WebSocket support is critical for low-latency bidirectional streaming. Full Azure Speech vs Amazon Polly verdict →

Both tools support streaming output and instant voice cloning, which are baseline requirements for voice agents.
See pricingWebsite →

Both tools support streaming output and instant voice cloning, which are baseline requirements for voice agents. Full Fish Speech vs CosyVoice verdict →

For voice agents, latency is the deciding factor.
See pricingTry Murf API →

For voice agents, latency is the deciding factor. Full Murf API vs Azure Speech verdict →

For real-time voice agents, low latency and websocket streaming are critical. Full Murf API vs Speechify API verdict →

For voice agents, latency is the decisive factor.

For voice agents, latency is the decisive factor. Full Deepgram Aura-2 vs ElevenLabs verdict →

For voice agents, latency is decisive.

For voice agents, latency is decisive. Full Fish Audio vs MiniMax Speech verdict →

For voice agents, latency is decisive.
See pricingTry Rime →

For voice agents, latency is decisive. Full Rime vs Inworld TTS verdict →

For voice agents, low latency and real-time bidirectional streaming are critical.

For voice agents, low latency and real-time bidirectional streaming are critical. Full OpenAI TTS vs Voxtral TTS verdict →

For realtime voice agents, websocket support is critical for low-latency bidirectional audio. Full OpenAI TTS vs Google Cloud TTS verdict →

For voice agents, two practical factors stand out.

For voice agents, two practical factors stand out. Full Google Cloud TTS vs Amazon Polly verdict →

Cloud-utility TTS at commodity prices.

Cloud-utility TTS at commodity prices. No won verdicts for this use case yet; it ranks on ties and near-misses.

Full-stack open TTS from Alibaba's speech team: 9 languages plus 18+ Chinese dialects, 150 ms streaming latency claim, instruction control of emotion/dialect/speed, and training + deployment scripts under Apache-2.0.
See pricingWebsite →

Full-stack open TTS from Alibaba's speech team: 9 languages plus 18+ Chinese dialects, 150 ms streaming latency claim, instruction control of emotion/dialect/speed, and training + deployment scripts under Apache-2.0. No won verdicts for this use case yet; it ranks on ties and near-misses.

Best-known open model for scripted two-speaker dialogue rather than narration - Apache-2.0, English-only, GPU-oriented.
See pricingWebsite →

Best-known open model for scripted two-speaker dialogue rather than narration - Apache-2.0, English-only, GPU-oriented. No won verdicts for this use case yet; it ranks on ties and near-misses.

Cost-leader realtime TTS for voice agents and games; hybrid pay-as-you-go plus monthly credit plans that lower the per-1M-character rate as commitment grows; SOC 2 Type II with zero-data-retention posture.

Cost-leader realtime TTS for voice agents and games; hybrid pay-as-you-go plus monthly credit plans that lower the per-1M-character rate as commitment grows; SOC 2 Type II with zero-data-retention posture. No won verdicts for this use case yet; it ranks on ties and near-misses.

Speed-and-affordability challenger: a single fast model (Blizzard 2) with streaming-first APIs, unlimited voice clones on every tier, and flat subscription pricing with per-1K overage.
From $10/moTry LMNT →

Speed-and-affordability challenger: a single fast model (Blizzard 2) with streaming-first APIs, unlimited voice clones on every tier, and flat subscription pricing with per-1K overage. No won verdicts for this use case yet; it ranks on ties and near-misses.

Multilingual cloning-first TTS with aggressive pricing.

Multilingual cloning-first TTS with aggressive pricing. No won verdicts for this use case yet; it ranks on ties and near-misses.

Developer platform spun out of the Speechify brand: transparent tiered pricing ($10-$499/mo plus per-1M overage), streaming-native Simba 3.2, and a bundled voice-agents product with flat per-minute rates.

Developer platform spun out of the Speechify brand: transparent tiered pricing ($10-$499/mo plus per-1M overage), streaming-native Simba 3.2, and a bundled voice-agents product with flat per-minute rates. No won verdicts for this use case yet; it ranks on ties and near-misses.

Open-weights-friendly voice cloning TTS from a frontier AI lab.

Open-weights-friendly voice cloning TTS from a frontier AI lab. No won verdicts for this use case yet; it ranks on ties and near-misses.

More Text-to-speech APIs buyer guides