# Best Text-to-speech APIs for Voice Agents (2026)

> The best Text-to-speech APIs platforms for voice agents: Cartesia leads, for realtime speech for phone agents and voice bots, where latency and streaming.

For voice agents, **Cartesia** is our pick (from $5/mo): For real-time voice agents, websocket streaming is critical for low-latency bidirectional conversation. Realtime speech for phone agents and voice bots, where latency and streaming decide whether the conversation feels human. Below is the full ranking and the tradeoffs, or read [how we score](https://www.versusref.com/methodology/).

## What matters for voice agents

Weight ×5 = decisive, ×1 = relevant.

| Fact | Weight | Cartesia | ElevenLabs | Azure Speech | Fish Speech | Murf API |
| --- | --- | --- | --- | --- | --- | --- |
| TTFB latency (vendor-claimed) | ×5 | ~90 ms (Jul 20) | ~75 ms (Jul 20) | n/a | n/a | ~130 ms (Jul 20) |
| Streaming audio output | ×5 | ✓  Yes (Jul 20) | ✓  Yes (Jul 20) | ✓  Yes (Jul 20) | ✓  Yes (Jul 20) | ✓  Yes (Jul 20) |
| Realtime websocket API | ×4 | ✓  Yes (Jul 20) | ✓  Yes (Jul 20) | ✓  Yes (Jul 20) | n/a | ✓  Yes (Jul 20) |
| Price per 1M characters (flagship model) | ×3 | n/a | 100 $/1M chars (Jul 20) | 22 $/1M chars (Jul 20) | n/a | 30 $/1M chars (Jul 20) |
| Concurrency on base plan | ×3 | Free: 2 (Jul 20) | Free: 2 (Multilingual v2) / 4 (Flash) (Jul 20) | F0: 20 transactions per 60 seconds (Jul 20) | n/a | 5 (Jul 20) |

- ×5 **TTFB latency (vendor-claimed):** Time to first audio byte is the single biggest driver of how responsive a voice agent feels.
- ×5 **Streaming audio output:** Agents must start speaking before the full reply is synthesized; non-streaming APIs are a hard stop.
- ×4 **Realtime websocket API:** A persistent websocket avoids per-request connection overhead in live conversations.
- ×3 **Price per 1M characters (flagship model):** Per-character rate compounds fast at call-center volumes.
- ×3 **Concurrency on base plan:** Concurrent call capacity on the entry plan decides when you are forced into a bigger contract.

## The ranking, tool by tool

| Rank | Tool | Verdict | Score | Price |
| --- | --- | --- | --- | --- |
| 1 | [Cartesia](https://www.versusref.com/tts/tools/cartesia/) | For real-time voice agents, websocket streaming is critical for low-latency bidirectional conversation. | 10 of 14 points · 7 matchups | From $5/mo |
| 2 | [ElevenLabs](https://www.versusref.com/tts/tools/elevenlabs/) | For real-time voice agents, the decisive feature is a WebSocket API for bidirectional streaming. | 11 of 20 points · 10 matchups | From $6/mo |
| 3 | [Azure Speech](https://www.versusref.com/tts/tools/azure-speech/) | Both tools offer real-time WebSocket APIs and streaming output, so latency infrastructure is comparable. | 5 of 10 points · 5 matchups | From $960/mo |
| 4 | [Fish Speech](https://www.versusref.com/tts/tools/fish-speech/) (OSS) | Both tools support streaming output and instant voice cloning, which are baseline requirements for voice agents. | 1 of 2 points · 1 matchup | See pricing |
| 5 | [Murf API](https://www.versusref.com/tts/tools/murf/) | For voice agents, latency is the deciding factor. | 2 of 6 points · 3 matchups | See pricing |
| 6 | [Deepgram Aura-2](https://www.versusref.com/tts/tools/deepgram-aura/) | For voice agents, latency is the decisive factor. | 1 of 4 points · 2 matchups | See pricing |
| 7 | [Fish Audio](https://www.versusref.com/tts/tools/fish-audio/) | For voice agents, latency is decisive. | 1 of 4 points · 2 matchups | See pricing |
| 8 | [Rime](https://www.versusref.com/tts/tools/rime/) | For voice agents, latency is decisive. | 1 of 4 points · 2 matchups | See pricing |
| 9 | [OpenAI TTS](https://www.versusref.com/tts/tools/openai-tts/) | For voice agents, low latency and real-time bidirectional streaming are critical. | 2 of 10 points · 5 matchups | See pricing |
| 10 | [Google Cloud TTS](https://www.versusref.com/tts/tools/google-tts/) | For voice agents, two practical factors stand out. | 1 of 8 points · 4 matchups | See pricing |
| 11 | [Amazon Polly](https://www.versusref.com/tts/tools/amazon-polly/) | Cloud-utility TTS at commodity prices. | 0 of 6 points · 3 matchups | See pricing |
| 12 | [CosyVoice](https://www.versusref.com/tts/tools/cosyvoice/) (OSS) | Full-stack open TTS from Alibaba's speech team: 9 languages plus 18+ Chinese dialects, 150 ms streaming latency claim, instruction control of emotion/dialect/speed, and training + deployment scripts under Apache-2.0. | 0 of 2 points · 1 matchup | See pricing |
| 13 | [Dia / Dia2](https://www.versusref.com/tts/tools/dia/) (OSS) | Best-known open model for scripted two-speaker dialogue rather than narration - Apache-2.0, English-only, GPU-oriented. | 0 of 2 points · 1 matchup | See pricing |
| 14 | [Inworld TTS](https://www.versusref.com/tts/tools/inworld-tts/) | Cost-leader realtime TTS for voice agents and games; hybrid pay-as-you-go plus monthly credit plans that lower the per-1M-character rate as commitment grows; SOC 2 Type II with zero-data-retention posture. | 0 of 4 points · 2 matchups | From $25/mo |
| 15 | [LMNT](https://www.versusref.com/tts/tools/lmnt/) | Speed-and-affordability challenger: a single fast model (Blizzard 2) with streaming-first APIs, unlimited voice clones on every tier, and flat subscription pricing with per-1K overage. | 0 of 2 points · 1 matchup | From $10/mo |
| 16 | [MiniMax Speech](https://www.versusref.com/tts/tools/minimax-speech/) | Multilingual cloning-first TTS with aggressive pricing. | 0 of 2 points · 1 matchup | From $5/mo |
| 17 | [Speechify API](https://www.versusref.com/tts/tools/speechify-api/) | Developer platform spun out of the Speechify brand: transparent tiered pricing ($10-$499/mo plus per-1M overage), streaming-native Simba 3.2, and a bundled voice-agents product with flat per-minute rates. | 0 of 2 points · 1 matchup | From $10/mo |
| 18 | [Voxtral TTS](https://www.versusref.com/tts/tools/voxtral-tts/) | Open-weights-friendly voice cloning TTS from a frontier AI lab. | 0 of 6 points · 3 matchups | See pricing |

### 1. Cartesia

For real-time voice agents, websocket streaming is critical for low-latency bidirectional conversation. [Full Cartesia vs Voxtral TTS verdict](https://www.versusref.com/tts/cartesia-vs-voxtral-tts/)

For voice agents, latency is decisive. [Full Cartesia vs Rime verdict](https://www.versusref.com/tts/cartesia-vs-rime/)

For voice agents, latency is the decisive factor. [Full Cartesia vs OpenAI TTS verdict](https://www.versusref.com/tts/cartesia-vs-openai-tts/)

For voice agents, latency is the decisive factor. [Full Cartesia vs LMNT verdict](https://www.versusref.com/tts/cartesia-vs-lmnt/)

For voice agents, latency is decisive. [Full Cartesia vs Inworld TTS verdict](https://www.versusref.com/tts/cartesia-vs-inworld-tts/)

For voice agents, latency is decisive. [Full Cartesia vs ElevenLabs verdict](https://www.versusref.com/tts/cartesia-vs-elevenlabs/)

For voice agents, latency is the critical axis. [Full Cartesia vs Deepgram Aura-2 verdict](https://www.versusref.com/tts/cartesia-vs-deepgram-aura/)

### 2. ElevenLabs

For real-time voice agents, the decisive feature is a WebSocket API for bidirectional streaming. [Full ElevenLabs vs Voxtral TTS verdict](https://www.versusref.com/tts/elevenlabs-vs-voxtral-tts/)

Both tools support real-time WebSocket API and streaming output, so baseline infrastructure is equal. [Full ElevenLabs vs OpenAI TTS verdict](https://www.versusref.com/tts/elevenlabs-vs-openai-tts/)

For real-time voice agents, latency is the decisive factor. [Full ElevenLabs vs Murf API verdict](https://www.versusref.com/tts/elevenlabs-vs-murf/)

For voice agents, a realtime WebSocket API is critical. [Full ElevenLabs vs Google Cloud TTS verdict](https://www.versusref.com/tts/elevenlabs-vs-google-tts/)

For real-time voice agents, latency is decisive. [Full ElevenLabs vs Fish Audio verdict](https://www.versusref.com/tts/elevenlabs-vs-fish-audio/)

ElevenLabs has a vendor-claimed TTFB of 75 ms and a realtime WebSocket API (fact 0c0932fb), both critical for natural-feeling voice agent conversations. [Full ElevenLabs vs Dia / Dia2 verdict](https://www.versusref.com/tts/dia-vs-elevenlabs/)

Both tools offer real-time WebSocket APIs and streaming output, so core infrastructure is equal. [Full ElevenLabs vs Azure Speech verdict](https://www.versusref.com/tts/azure-speech-vs-elevenlabs/)

For voice agents, realtime responsiveness is decisive. [Full ElevenLabs vs Amazon Polly verdict](https://www.versusref.com/tts/amazon-polly-vs-elevenlabs/)

### 3. Azure Speech

Both tools offer real-time WebSocket APIs and streaming output, so latency infrastructure is comparable. [Full Azure Speech vs OpenAI TTS verdict](https://www.versusref.com/tts/azure-speech-vs-openai-tts/)

For voice agents, real-time responsiveness is critical. [Full Azure Speech vs Google Cloud TTS verdict](https://www.versusref.com/tts/azure-speech-vs-google-tts/)

For real-time voice agent use cases, WebSocket support is critical for low-latency bidirectional streaming. [Full Azure Speech vs Amazon Polly verdict](https://www.versusref.com/tts/amazon-polly-vs-azure-speech/)

### 4. Fish Speech

Both tools support streaming output and instant voice cloning, which are baseline requirements for voice agents. [Full Fish Speech vs CosyVoice verdict](https://www.versusref.com/tts/cosyvoice-vs-fish-speech/)

### 5. Murf API

For voice agents, latency is the deciding factor. [Full Murf API vs Azure Speech verdict](https://www.versusref.com/tts/azure-speech-vs-murf/)

For real-time voice agents, low latency and websocket streaming are critical. [Full Murf API vs Speechify API verdict](https://www.versusref.com/tts/murf-vs-speechify-api/)

### 6. Deepgram Aura-2

For voice agents, latency is the decisive factor. [Full Deepgram Aura-2 vs ElevenLabs verdict](https://www.versusref.com/tts/deepgram-aura-vs-elevenlabs/)

### 7. Fish Audio

For voice agents, latency is decisive. [Full Fish Audio vs MiniMax Speech verdict](https://www.versusref.com/tts/fish-audio-vs-minimax-speech/)

### 8. Rime

For voice agents, latency is decisive. [Full Rime vs Inworld TTS verdict](https://www.versusref.com/tts/inworld-tts-vs-rime/)

### 9. OpenAI TTS

For voice agents, low latency and real-time bidirectional streaming are critical. [Full OpenAI TTS vs Voxtral TTS verdict](https://www.versusref.com/tts/openai-tts-vs-voxtral-tts/)

For realtime voice agents, websocket support is critical for low-latency bidirectional audio. [Full OpenAI TTS vs Google Cloud TTS verdict](https://www.versusref.com/tts/google-tts-vs-openai-tts/)

### 10. Google Cloud TTS

For voice agents, two practical factors stand out. [Full Google Cloud TTS vs Amazon Polly verdict](https://www.versusref.com/tts/amazon-polly-vs-google-tts/)

### 11. Amazon Polly

Cloud-utility TTS at commodity prices. No won verdicts for this use case yet; it ranks on ties and near-misses.

### 12. CosyVoice

Full-stack open TTS from Alibaba's speech team: 9 languages plus 18+ Chinese dialects, 150 ms streaming latency claim, instruction control of emotion/dialect/speed, and training + deployment scripts under Apache-2.0. No won verdicts for this use case yet; it ranks on ties and near-misses.

### 13. Dia / Dia2

Best-known open model for scripted two-speaker dialogue rather than narration - Apache-2.0, English-only, GPU-oriented. No won verdicts for this use case yet; it ranks on ties and near-misses.

### 14. Inworld TTS

Cost-leader realtime TTS for voice agents and games; hybrid pay-as-you-go plus monthly credit plans that lower the per-1M-character rate as commitment grows; SOC 2 Type II with zero-data-retention posture. No won verdicts for this use case yet; it ranks on ties and near-misses.

### 15. LMNT

Speed-and-affordability challenger: a single fast model (Blizzard 2) with streaming-first APIs, unlimited voice clones on every tier, and flat subscription pricing with per-1K overage. No won verdicts for this use case yet; it ranks on ties and near-misses.

### 16. MiniMax Speech

Multilingual cloning-first TTS with aggressive pricing. No won verdicts for this use case yet; it ranks on ties and near-misses.

### 17. Speechify API

Developer platform spun out of the Speechify brand: transparent tiered pricing ($10-$499/mo plus per-1M overage), streaming-native Simba 3.2, and a bundled voice-agents product with flat per-minute rates. No won verdicts for this use case yet; it ranks on ties and near-misses.

### 18. Voxtral TTS

Open-weights-friendly voice cloning TTS from a frontier AI lab. No won verdicts for this use case yet; it ranks on ties and near-misses.

Source: https://www.versusref.com/tts/best/voice-agents/
