# Text-to-speech APIs platforms

> Compare 52 Text-to-speech APIs platforms on 742 verified facts, from ElevenLabs vs OpenAI TTS to the rest. Real production cost and per-use-case verdicts.

52 platforms tracked · 742 facts verified · pricing re-checked weekly. APIs and platforms that turn text into speech: real per-minute cost, latency, cloning rights, and a verdict per use case.

Last verified Jul 20, 2026

## Text-to-speech APIs platforms we track

| Tool | Positioning | Price | Rating |
| --- | --- | --- | --- |
| [Amazon Polly](https://www.versusref.com/tts/tools/amazon-polly/) | Cloud-utility TTS at commodity prices | See pricing |  |
| [Azure Speech](https://www.versusref.com/tts/tools/azure-speech/) | Enterprise hyperscaler TTS with custom-voice depth | From $960/mo |  |
| [Cartesia](https://www.versusref.com/tts/tools/cartesia/) | Lowest-latency TTS for real-time voice agents | From $5/mo |  |
| [Deepgram Aura-2](https://www.versusref.com/tts/tools/deepgram-aura/) | Enterprise real-time voice-agent TTS | See pricing |  |
| [ElevenLabs](https://www.versusref.com/tts/tools/elevenlabs/) | Premium AI voice platform for creators and developers | From $6/mo |  |
| [Google Cloud TTS](https://www.versusref.com/tts/tools/google-tts/) | Hyperscaler TTS with the broadest voice/language catalog | See pricing |  |
| [MiniMax Speech](https://www.versusref.com/tts/tools/minimax-speech/) | Multilingual cloning-first TTS with aggressive pricing | From $5/mo |  |
| [OpenAI TTS](https://www.versusref.com/tts/tools/openai-tts/) | Simple usage-based TTS inside a general AI platform | See pricing |  |
| [Voxtral TTS](https://www.versusref.com/tts/tools/voxtral-tts/) | Open-weights-friendly voice cloning TTS from a frontier AI lab | See pricing |  |
| [CAMB.AI](https://www.versusref.com/tts/tools/camb-ai/) | Localization-first TTS: the MARS 8 family (flash/pro/instruct variants) plus dubbing and translated-TTS pipelines, credit-based plans from $5/mo, aimed at media, sports, and content going multilingual. | From $5/mo |  |
| [Fish Audio](https://www.versusref.com/tts/tools/fish-audio/) | Developer-first hosted API from the team behind the open-source fish-speech models; simple prepaid pay-as-you-go billing, 80+ language coverage, and a free fair-use model tier (s2.1-pro-free). Facts here cover the hosted API, not the OSS model. | See pricing |  |
| [Hume Octave TTS](https://www.versusref.com/tts/tools/hume-octave/) | Expressiveness-first TTS (Octave understands the meaning of the text it speaks); subscription plans gate commercial use, with Octave 2 (preview) adding ~100ms latency and 11 languages for realtime use. | From $3/mo |  |
| [Inworld TTS](https://www.versusref.com/tts/tools/inworld-tts/) | Cost-leader realtime TTS for voice agents and games; hybrid pay-as-you-go plus monthly credit plans that lower the per-1M-character rate as commitment grows; SOC 2 Type II with zero-data-retention posture. | From $25/mo |  |
| [LMNT](https://www.versusref.com/tts/tools/lmnt/) | Speed-and-affordability challenger: a single fast model (Blizzard 2) with streaming-first APIs, unlimited voice clones on every tier, and flat subscription pricing with per-1K overage. | From $10/mo |  |
| [Murf API](https://www.versusref.com/tts/tools/murf/) | API arm of the Murf studio platform: 150+ voices in 35 languages, SSML support, word timestamps, and Falcon 2 aimed at high-concurrency voice agents at $0.01/1K characters. Note: Murf Studio subscription plans (murf.ai/pricing) are a separate product from API pay-as-you-go pricing and API characters are purchased separately. | See pricing |  |
| [Resemble AI](https://www.versusref.com/tts/tools/resemble/) | Security-first enterprise play: generation plus detection/verification in one platform, pay-as-you-go Flex credits, on-prem option, and the MIT-licensed open-source Chatterbox model family. | See pricing |  |
| [Rime](https://www.versusref.com/tts/tools/rime/) | Enterprise conversational TTS (IVR, contact centers, voice agents) emphasizing ultra-low latency models (Coda, Mist, Arcana) and self-hosted deployment; usage-based pricing with a single published rate. | See pricing |  |
| [Speechify API](https://www.versusref.com/tts/tools/speechify-api/) | Developer platform spun out of the Speechify brand: transparent tiered pricing ($10-$499/mo plus per-1M overage), streaming-native Simba 3.2, and a bundled voice-agents product with flat per-minute rates. | From $10/mo |  |
| [WellSaid](https://www.versusref.com/tts/tools/wellsaid/) | Enterprise voiceover specialist: polished Studio product for L&D/marketing narration with an API on the side; API pricing is contact-sales, and compliance (SOC 2 Type II, GDPR) and ethical voice sourcing are the pitch. | From $10/mo |  |
| [Chatterbox](https://www.versusref.com/tts/tools/chatterbox/) (OSS) | Production-minded open TTS from a commercial voice company - MIT license, built-in watermarking, and a fast Turbo variant, with Resemble's paid API as the scale-up path. | See pricing |  |
| [ChatTTS](https://www.versusref.com/tts/tools/chattts/) (OSS) | Optimized for natural dialogue-style speech for LLM assistants; the licensing combination (AGPLv3+ code, CC BY-NC 4.0 weights, research/education only) rules out most commercial SaaS use without a separate deal. | See pricing |  |
| [CosyVoice](https://www.versusref.com/tts/tools/cosyvoice/) (OSS) | Full-stack open TTS from Alibaba's speech team: 9 languages plus 18+ Chinese dialects, 150 ms streaming latency claim, instruction control of emotion/dialect/speed, and training + deployment scripts under Apache-2.0. | See pricing |  |
| [Dia / Dia2](https://www.versusref.com/tts/tools/dia/) (OSS) | Best-known open model for scripted two-speaker dialogue rather than narration - Apache-2.0, English-only, GPU-oriented. | See pricing |  |
| [F5-TTS](https://www.versusref.com/tts/tools/f5-tts/) (OSS) | The go-to research-grade voice-cloning model - actively maintained and broadly ported, with the classic code-vs-weights license split: commercial products must retrain or license around the CC-BY-NC checkpoints. | See pricing |  |
| [Fish Speech](https://www.versusref.com/tts/tools/fish-speech/) (OSS) | Top-tier expressive multilingual open-weights TTS whose license moved from permissive to research/non-commercial; commercial use requires a license from Fish Audio or their hosted API. | See pricing |  |
| [GPT-SoVITS](https://www.versusref.com/tts/tools/gpt-sovits/) (OSS) | The de facto community standard for DIY voice cloning (60k GitHub stars), with a full WebUI covering dataset prep, ASR, training, and inference across five languages; MIT-licensed. | See pricing |  |
| [Gradium](https://www.versusref.com/tts/tools/gradium/) | Real-time voice-agent infrastructure play: WebSocket-first streaming TTS in 5 European languages with instant cloning, on-device models (Phonon), and credit-based pricing. Active and well-funded (site announced funding extension to $100M, July 2026). | From $13/mo |  |
| [Grok TTS](https://www.versusref.com/tts/tools/grok-tts/) | Usage-priced hosted TTS from xAI, part of the Grok Voice stack (TTS, STT, and a speech-to-speech Voice Agent API), aimed at developers who want low-latency voice output alongside Grok models. | See pricing |  |
| [Higgs Audio](https://www.versusref.com/tts/tools/higgs-audio/) (OSS) | Strong emotional/expressive open TTS from a well-funded lab - but 'Apache-2.0' only covers the repo code; V2 weights carry a 100k-MAU community license and the newer V3 is research/non-commercial. | See pricing |  |
| [IndexTTS-2](https://www.versusref.com/tts/tools/index-tts/) (OSS) | Emotion-controllable zero-shot voice cloning for production use, but under a custom Bilibili license (not OSI-approved) with scale caps and AI-training restrictions. | See pricing |  |
| [KittenTTS](https://www.versusref.com/tts/tools/kitten-tts/) (OSS) | The lightweight/edge option: ONNX inference, sub-100 MB downloads, 8 built-in voices - trades voice cloning and multilinguality for footprint. | See pricing |  |
| [Kokoro](https://www.versusref.com/tts/tools/kokoro/) (OSS) | The lightweight quality-per-parameter champion of open TTS: Apache-2.0, easy to run anywhere, no voice cloning by design. | See pricing |  |
| [Kyutai TTS](https://www.versusref.com/tts/tools/kyutai-tts/) (OSS) | Research-lab open TTS optimized for real-time streaming (Delayed Streams Modeling); Pocket TTS targets on-device/CPU deployment while the larger DSM TTS targets production streaming servers (Rust backend). | See pricing |  |
| [Maya1](https://www.versusref.com/tts/tools/maya1/) (OSS) | Apache 2.0 expressive English TTS you can run on a single 16GB+ GPU; stands out for natural-language voice design and 20+ inline emotion tags rather than audio-sample cloning, with real-time streaming via vLLM. | See pricing |  |
| [MegaTTS3](https://www.versusref.com/tts/tools/megatts3/) (OSS) | Research-grade Apache-2.0 TTS whose practical cloning is gated: the WaveVAE encoder is not released, so users must submit audio to ByteDance channels to obtain pre-extracted speaker latents (.npy) for cloning. | See pricing |  |
| [Narakeet](https://www.versusref.com/tts/tools/narakeet/) | Indie batch-content workhorse: aggregates a very broad voice/language catalog for voiceovers, audiobooks and video narration, priced per output minute (prepaid packs, no subscription) rather than per character; not aimed at real-time agent use. | See pricing |  |
| [Neuphonic](https://www.versusref.com/tts/tools/neuphonic/) | Hybrid hosted + open-weights play: SSE/WebSocket streaming TTS API at app.neuphonic.com, and tiny CPU-only on-device models (NeuTTS-Air ~360M Apache-2.0, NeuTTS-Nano ~120M) in GGUF for phones/Raspberry Pi - privacy/edge-deployment angle. NOTE: the site's pricing page returned 404 at verification time; hosted-plan pricing treated as not published. | See pricing |  |
| [NVIDIA Magpie TTS](https://www.versusref.com/tts/tools/nvidia-magpie/) (OSS) | A small, GPU-efficient 9-language TTS checkpoint for teams already in the NVIDIA NeMo/Riva ecosystem; commercially usable open weights, but zero-shot voice cloning was removed from the open release and it caps generations at about 20 seconds. | See pricing |  |
| [Orpheus TTS](https://www.versusref.com/tts/tools/orpheus/) (OSS) | The 'LLM-as-TTS' approach under Apache-2.0 - expressive and clonable, but repo activity has slowed since mid-2025. | See pricing |  |
| [OuteTTS](https://www.versusref.com/tts/tools/outetts/) (OSS) | The llama.cpp-native option: runs via GGUF on CUDA/ROCm/Vulkan/Metal and even in the browser (Transformers.js). License splits by model: the 0.6B (Qwen3-based) is Apache-2.0, the 1B Llama-based flagship is CC-BY-NC-SA-4.0 (non-commercial). | See pricing |  |
| [Piper](https://www.versusref.com/tts/tools/piper/) (OSS) | The pragmatic embedded/self-hosted choice: no cloning or frills, just quick offline speech in many languages - note the license change from MIT (archived rhasspy/piper) to GPL-3.0 in the successor repo. | See pricing |  |
| [Qwen3-TTS](https://www.versusref.com/tts/tools/qwen3-tts/) (OSS) | Genuinely open weights (confirmed - not API-only since the Jan 2026 release): Apache-2.0 checkpoints on Hugging Face with an Alibaba Cloud DashScope API for hosted use. | See pricing |  |
| [Sarvam AI Bulbul](https://www.versusref.com/tts/tools/sarvam-bulbul/) | The go-to hosted TTS for Indian-language products (Hindi, Tamil, Telugu, Bengali and more) with REST, HTTP streaming and WebSocket APIs and prepaid INR pricing; not aimed at global multilingual coverage. | See pricing |  |
| [Smallest.ai Waves](https://www.versusref.com/tts/tools/smallest-ai/) | Speed- and price-led challenger from an India-focused voice AI startup; Waves is the speech-model API layer (Lightning TTS, Pulse STT), sold pay-as-you-go with $10 free credits, with HIPAA/SOC2/on-prem reserved for the Enterprise plan. | See pricing |  |
| [Speechmatics TTS](https://www.versusref.com/tts/tools/speechmatics-tts/) | Budget low-latency English TTS for voice agents; priced at a fraction of premium voice APIs, but currently English-only with a small voice set and no SSML or voice cloning. | See pricing |  |
| [Step-Audio](https://www.versusref.com/tts/tools/step-audio/) (OSS) | Differentiates on iterative audio *editing* (emotion, style, breathing, laughter, sighs, polyphone pinyin control) rather than plain synthesis; Apache-2.0 with training code (SFT/DPO/GRPO) and vLLM support. | See pricing |  |
| [Supertone API](https://www.versusref.com/tts/tools/supertone/) | Expressive character voices (games, content, Korean/Japanese/English markets first, now 31 languages via Supertonic 3) on cheap credit-based subscriptions from $2.99/month; uniquely pairs the hosted API with the open OpenRAIL-M Supertonic model for on-device synthesis. | From $2.99/mo |  |
| [Typecast](https://www.versusref.com/tts/tools/typecast/) | Expressiveness play from a Korean AI-actor studio: strongest on emotion prompts/presets and character voices for content and conversational AI; separate consumer studio subscription ($8.99+) and developer API plans (Free/Lite/Plus). | From $15/mo |  |
| [Unreal Speech](https://www.versusref.com/tts/tools/unreal-speech/) | Pure price play: markets itself as up to 11x cheaper than ElevenLabs with a simple 3-endpoint API (stream/speech/synthesisTasks); smaller voice/language catalog and no voice cloning documented. | From $49/mo |  |
| [VibeVoice](https://www.versusref.com/tts/tools/vibevoice/) (OSS) | The open long-form/multi-speaker specialist - MIT weights, but Microsoft pulled the TTS code from the repo in Sept 2025 and frames the models as research-only. | See pricing |  |
| [Voxtral (open weights)](https://www.versusref.com/tts/tools/voxtral-open/) (OSS) | A frontier-lab open-weight TTS you can run on a single 16GB GPU; the CC BY-NC license makes it evaluation/research-only, with Mistral's paid API as the commercial route. | See pricing |  |
| [XTTS v2 (Coqui)](https://www.versusref.com/tts/tools/xtts/) (OSS) | The legacy standard for open voice cloning: dormant upstream since Coqui's Jan 2024 shutdown (community fork idiap/coqui-ai-TTS carries maintenance), and CPML weights bar commercial use. | See pricing |  |

## Text-to-speech APIs pricing compared (2026)

Effective monthly bill at 2M chars/mo (product feature), computed from each vendor's published rates - cheapest first. Every tool page carries the full three-tier table.

| Tool | Published rate | Bill at 2M chars/mo | Verified |
| --- | --- | --- | --- |
| [Supertone API](https://www.versusref.com/tts/tools/supertone/) | Cheapest paid plan $2.99/mo | $2.99 | [Jul 20](https://www.supertone.ai/en/api) |
| [CAMB.AI](https://www.versusref.com/tts/tools/camb-ai/) | Cheapest paid plan $5/mo | $5 | [Jul 20](https://www.camb.ai/pricing) |
| [WellSaid](https://www.versusref.com/tts/tools/wellsaid/) | Cheapest paid plan $10/mo | $10 | [Jul 20](https://www.wellsaid.io/pricing) |
| [Speechify API](https://www.versusref.com/tts/tools/speechify-api/) | Usage-priced at $10 per 1M characters (≈ $0.0095 per audio-minute) | $20 | [Jul 20](https://speechify.ai/pricing) |
| [Speechmatics TTS](https://www.versusref.com/tts/tools/speechmatics-tts/) | Usage-priced at $11 per 1M characters (≈ $0.0105 per audio-minute) | $22 | [Jul 20](https://www.speechmatics.com/text-to-speech) |
| [Fish Audio](https://www.versusref.com/tts/tools/fish-audio/) | Usage-priced at $15 per 1M characters (≈ $0.0143 per audio-minute) | $30 | [Jul 20](https://docs.fish.audio/developer-guide/models-pricing/pricing-and-rate-limits) |
| [Grok TTS](https://www.versusref.com/tts/tools/grok-tts/) | Usage-priced at $15 per 1M characters (≈ $0.0143 per audio-minute) | $30 | [Jul 20](https://docs.x.ai/docs/models) |
| [Voxtral TTS](https://www.versusref.com/tts/tools/voxtral-tts/) | Usage-priced at $16 per 1M characters (≈ $0.0152 per audio-minute) | $32 | [Jul 20](https://mistral.ai/news/voxtral-tts) |
| [Azure Speech](https://www.versusref.com/tts/tools/azure-speech/) | Usage-priced at $22 per 1M characters (≈ $0.0209 per audio-minute) | $44 | [Jul 20](https://azure.microsoft.com/en-us/pricing/details/cognitive-services/speech-services/) |
| [Unreal Speech](https://www.versusref.com/tts/tools/unreal-speech/) | Cheapest paid plan $49/mo with 3,000,000 characters included | $49 | [Jul 20](https://unrealspeech.com/pricing) |
| [Inworld TTS](https://www.versusref.com/tts/tools/inworld-tts/) | Usage-priced at $25 per 1M characters (≈ $0.0238 per audio-minute) | $50 | [Jul 20](https://inworld.ai/pricing) |
| [Amazon Polly](https://www.versusref.com/tts/tools/amazon-polly/) | Usage-priced at $30 per 1M characters (≈ $0.0285 per audio-minute) | $60 | [Jul 20](https://aws.amazon.com/polly/pricing/) |
| [Deepgram Aura-2](https://www.versusref.com/tts/tools/deepgram-aura/) | Usage-priced at $30 per 1M characters (≈ $0.0285 per audio-minute) | $60 | [Jul 20](https://deepgram.com/pricing) |
| [Google Cloud TTS](https://www.versusref.com/tts/tools/google-tts/) | Usage-priced at $30 per 1M characters (≈ $0.0285 per audio-minute) | $60 | [Jul 20](https://cloud.google.com/text-to-speech/pricing) |
| [Murf API](https://www.versusref.com/tts/tools/murf/) | Usage-priced at $30 per 1M characters (≈ $0.0285 per audio-minute) | $60 | [Jul 20](https://help.murf.ai/murf-api-plans-and-limits) |
| [OpenAI TTS](https://www.versusref.com/tts/tools/openai-tts/) | Usage-priced at $30 per 1M characters (≈ $0.0285 per audio-minute) | $60 | [Jul 20](https://developers.openai.com/api/docs/pricing) |
| [LMNT](https://www.versusref.com/tts/tools/lmnt/) | Cheapest paid plan $10/mo with 200,000 characters included | $100 | [Jul 20](https://www.lmnt.com/pricing/) |
| [Rime](https://www.versusref.com/tts/tools/rime/) | Usage-priced at $50 per 1M characters (≈ $0.0475 per audio-minute) | $100 | [Jul 20](https://rime.ai/pricing) |
| [Typecast](https://www.versusref.com/tts/tools/typecast/) | Usage-priced at $70 per 1M characters (≈ $0.0665 per audio-minute) | $140 | [Jul 20](https://typecast.ai/developers/api) |
| [ElevenLabs](https://www.versusref.com/tts/tools/elevenlabs/) | Usage-priced at $100 per 1M characters (≈ $0.095 per audio-minute) | $200 | [Jul 20](https://elevenlabs.io/pricing/api) |
| [MiniMax Speech](https://www.versusref.com/tts/tools/minimax-speech/) | Usage-priced at $100 per 1M characters (≈ $0.095 per audio-minute) | $200 | [Jul 20](https://platform.minimax.io/docs/guides/pricing-paygo.md) |
| [Cartesia](https://www.versusref.com/tts/tools/cartesia/) | Cheapest paid plan $5/mo with 100,000 characters included | Higher plan | [Jul 20](https://cartesia.ai/pricing) |
| [Gradium](https://www.versusref.com/tts/tools/gradium/) | Cheapest paid plan $13/mo with 225,000 characters included | Higher plan | [Jul 20](https://gradium.ai/pricing) |
| [Hume Octave TTS](https://www.versusref.com/tts/tools/hume-octave/) | Cheapest paid plan $3/mo with 30,000 characters included | Higher plan | [Jul 20](https://www.hume.ai/pricing) |

## Popular comparisons

- [ElevenLabs vs OpenAI TTS](https://www.versusref.com/tts/elevenlabs-vs-openai-tts/)
- [Cartesia vs ElevenLabs](https://www.versusref.com/tts/cartesia-vs-elevenlabs/)
- [ElevenLabs vs Voxtral TTS](https://www.versusref.com/tts/elevenlabs-vs-voxtral-tts/)
- [ElevenLabs vs Murf API](https://www.versusref.com/tts/elevenlabs-vs-murf/)
- [Amazon Polly vs ElevenLabs](https://www.versusref.com/tts/amazon-polly-vs-elevenlabs/)
- [ElevenLabs vs Google Cloud TTS](https://www.versusref.com/tts/elevenlabs-vs-google-tts/)
- [Azure Speech vs ElevenLabs](https://www.versusref.com/tts/azure-speech-vs-elevenlabs/)
- [ElevenLabs vs Fish Audio](https://www.versusref.com/tts/elevenlabs-vs-fish-audio/)
- [Deepgram Aura-2 vs ElevenLabs](https://www.versusref.com/tts/deepgram-aura-vs-elevenlabs/)
- [Cartesia vs OpenAI TTS](https://www.versusref.com/tts/cartesia-vs-openai-tts/)
- [Cartesia vs Rime](https://www.versusref.com/tts/cartesia-vs-rime/)
- [Cartesia vs Deepgram Aura-2](https://www.versusref.com/tts/cartesia-vs-deepgram-aura/)
- [Cartesia vs Voxtral TTS](https://www.versusref.com/tts/cartesia-vs-voxtral-tts/)
- [Google Cloud TTS vs OpenAI TTS](https://www.versusref.com/tts/google-tts-vs-openai-tts/)
- [Azure Speech vs OpenAI TTS](https://www.versusref.com/tts/azure-speech-vs-openai-tts/)
- [OpenAI TTS vs Voxtral TTS](https://www.versusref.com/tts/openai-tts-vs-voxtral-tts/)
- [Azure Speech vs Google Cloud TTS](https://www.versusref.com/tts/azure-speech-vs-google-tts/)
- [Amazon Polly vs Google Cloud TTS](https://www.versusref.com/tts/amazon-polly-vs-google-tts/)
- [Amazon Polly vs Azure Speech](https://www.versusref.com/tts/amazon-polly-vs-azure-speech/)
- [Murf API vs Speechify API](https://www.versusref.com/tts/murf-vs-speechify-api/)
- [Fish Audio vs MiniMax Speech](https://www.versusref.com/tts/fish-audio-vs-minimax-speech/)
- [Inworld TTS vs Rime](https://www.versusref.com/tts/inworld-tts-vs-rime/)
- [Cartesia vs Inworld TTS](https://www.versusref.com/tts/cartesia-vs-inworld-tts/)
- [Cartesia vs LMNT](https://www.versusref.com/tts/cartesia-vs-lmnt/)
- [CAMB.AI vs ElevenLabs](https://www.versusref.com/tts/camb-ai-vs-elevenlabs/)
- [Dia / Dia2 vs ElevenLabs](https://www.versusref.com/tts/dia-vs-elevenlabs/)
- [CosyVoice vs Fish Speech](https://www.versusref.com/tts/cosyvoice-vs-fish-speech/)
- [Azure Speech vs Murf API](https://www.versusref.com/tts/azure-speech-vs-murf/)

## Best for…

| Use case | Current pick | Guide |
| --- | --- | --- |
| Audiobooks | Azure Speech | [Best for audiobooks](https://www.versusref.com/tts/best/audiobooks/) |
| Content Creators | ElevenLabs | [Best for content creators](https://www.versusref.com/tts/best/content-creators/) |
| Developers | Cartesia | [Best for developers](https://www.versusref.com/tts/best/developers/) |
| Dubbing | Inworld TTS | [Best for dubbing](https://www.versusref.com/tts/best/dubbing-localization/) |
| Self-Hosted | Voxtral TTS | [Best for self-hosted](https://www.versusref.com/tts/best/self-hosted/) |
| Voice Agents | Cartesia | [Best for voice agents](https://www.versusref.com/tts/best/voice-agents/) |

Source: https://www.versusref.com/tts/
