vsref

Best Text-to-speech APIs for Developers (2026)

For developers, Cartesia is our pick (from $5/mo): For developers building speech products, Cartesia Sonic has several concrete advantages. Building speech into products: SDK quality, output flexibility, timestamps, and a usage-priced API you can meter. Below is the full ranking and the tradeoffs, or read how we score.

If you sign up through links on this page, vsref may earn a commission; programs exist on both sides of most comparisons, and commissions never change verdicts. How we make money

1Cartesia logoCartesiaWINNER
Lowest-latency TTS for real-time voice agents5 of 10 points · 5 matchups
Enterprise hyperscaler TTS with custom-voice depth3 of 6 points · 3 matchups
Hyperscaler TTS with the broadest voice/language catalog2 of 4 points · 2 matchups

What matters for developers

Weighted attribute comparison for Developers
FactCartesiaAzure SpeechGoogle Cloud TTSDeepgram Aura-2MiniMax Speech
Price per 1M characters (flagship model)×4n/a22 $/1M charsJul 2030 $/1M charsJul 2030 $/1M charsJul 20100 $/1M charsJul 20
Streaming audio output×4✓ YesJul 20✓ YesJul 20✓ YesJul 20✓ YesJul 20✓ YesJul 20
Official SDKs×4Python, JavaScript/TypeScriptJul 20Speech SDK: C#Jul 20C#Jul 20Official SDKs: Python, JavaScript/TypeScript, .NET, GoJul 20n/a
Word-level timestamps×3✓ YesJul 20✓ YesJul 20n/a✗ NoJul 20✓ YesJul 20
Output formats×3raw PCM (pcm_f32leJul 20MP3Jul 20MP3, LINEAR16 (WAV/PCM), OGG_OPUS, MULAW, ALAWJul 20mp3 (default), wav/linear16Jul 20mp3, pcm, flac, wav, pcmu_raw, pcmu_wav, opusJul 20
Swipe → to see every tool column.
×4 Price per 1M characters (flagship model): Usage pricing you can meter beats seat plans for product features.×4 Streaming audio output: Streaming synthesis unlocks responsive UX in chat and assistant features.×4 Official SDKs: Official SDKs in your stack cut integration time and maintenance risk.×3 Word-level timestamps: Word timing powers captions, karaoke highlighting, and precise UI sync.×3 Output formats: Native format/sample-rate options avoid a transcode step in your pipeline.

The ranking, tool by tool

1Cartesia logoCartesiaWINNER
For developers building speech products, Cartesia Sonic has several concrete advantages.

For developers building speech products, Cartesia Sonic has several concrete advantages. It supports a realtime WebSocket API while Mistral does not. Cartesia provides word-level timestamps and pronunciation dictionaries, both absent in Mistral. Cartesia supports 42 languages versus Mistral's 9. Both offer Python and TypeScript SDKs and streaming output. The WebSocket API and timestamps are especially critical for developer integrations that need low-latency interactivity and precise audio synchronization.

Cartesia offers official Python and JavaScript/TypeScript SDKs, while Rime provides only REST, WebSocket, and SSE APIs with no official SDKs. For developers building products, SDK availability meaningfully reduces integration friction. Both tools support word-level timestamps, streaming, and WebSocket APIs. Cartesia's vendor-claimed 90ms TTFB edges Rime's 120ms. Rime does offer 20 concurrent requests on its base plan versus Cartesia's 2 on the free tier, and Rime's 50-language coverage slightly exceeds Cartesia's 42. Even so, the SDK gap remains the decisive developer-experience factor here.

Both tools offer Python and TypeScript SDKs, WebSocket APIs, word-level timestamps, and streaming output. Cartesia edges ahead on latency at 90 ms versus LMNT at 150 ms, which matters for real-time product integrations. Cartesia supports 42 languages versus LMNT at 31, broadening addressable markets. Cartesia also adds pronunciation dictionaries and a credits-based pricing model suited to metered usage. LMNT counters with no concurrency limits on paid plans and an additional Go SDK, but the latency and language count advantages tip toward Cartesia for developer-focused builds.

Both tools share key developer features: WebSocket API, streaming, word timestamps, pronunciation dictionaries, and instant and professional voice cloning. Cartesia edges ahead on two concrete points. First, it offers both Python and JavaScript/TypeScript SDKs versus Inworld's Python-only SDK, directly serving more developer stacks. Second, Cartesia's TTFB is 90 ms versus Inworld's 200 ms, a meaningful latency advantage for interactive product experiences. Inworld supports more output formats and languages, but SDK breadth and latency tip the balance for product developers.

Both tools offer usage-based pricing and real-time websocket APIs, but Azure edges ahead on developer-focused features.

Both tools offer usage-based pricing and real-time websocket APIs, but Azure edges ahead on developer-focused features. Azure provides word-level timestamps, which OpenAI TTS lacks, along with SSML support and pronunciation dictionaries for precise speech control. Azure's flagship model is also priced at 22 dollars per 1M chars versus OpenAI's 30 dollars per 1M chars. OpenAI counters with more output formats, including Opus, AAC, FLAC, WAV, and PCM, compared to Azure's MP3, which matters for product flexibility. However, the timestamp gap is significant for many developer use cases, giving Azure the narrow win.

For developer metering, Azure charges 22 vs 100 per 1M chars for flagship voices and 15 vs 50 per 1M chars for fast voices, making it far cheaper to pass costs to users. Azure's pure usage pricing also avoids ElevenLabs' hybrid platform fees. Both offer word-level timestamps, SSML, streaming, and websocket APIs. Azure supports 100 languages vs 32 for ElevenLabs, broadening product reach. Azure's self-host option and HIPAA BAA without enterprise gating add deployment flexibility. ElevenLabs has a lower entry fee and a 75ms TTFB claim, which keeps the margin narrow.

Both tools have usage pricing, SSML, word timestamps, and good SDKs. Azure edges ahead on developer flexibility: it supports a realtime WebSocket API while Polly does not, enabling lower-latency interactive products. Azure also supports 100 languages vs Polly's 40, broadening addressable markets. Azure's flagship rate of $22/1M chars is lower than Polly's $30/1M chars. Polly's fast model is cheaper at $4 vs $15, but the WebSocket advantage is decisive for interactive developer use cases.

For metered usage pricing, Google charges 4 dollars per 1M characters (fast tier) versus ElevenLabs at 50 dollars per 1M characters, making Google far cheaper at scale.

For metered usage pricing, Google charges 4 dollars per 1M characters (fast tier) versus ElevenLabs at 50 dollars per 1M characters, making Google far cheaper at scale. Google also provides 4M free characters per month with commercial use allowed, while ElevenLabs restricts commercial use on its free tier. Google supports 1,000 requests per minute concurrency by default. ElevenLabs does offer advantages: a real-time WebSocket API (Google does not), word-level timestamps, and a 40,000 character per request limit versus Google's 5,000. Even so, the pricing difference is decisive for a metered use case.

Both tools share identical flagship and fast pricing at 30 and 4 dollars per 1M chars respectively, both usage-priced. Google edges ahead on developer flexibility: it supports 380 voices versus Amazon Polly's 100, handles 5000 chars per request versus Polly's 3000, and defaults to 1000 requests per minute versus Polly's 80 concurrent. Google also offers instant voice cloning, which Polly lacks. Polly counters with word-level timestamps and pronunciation dictionaries, useful for precise audio alignment. Overall, Google's higher concurrency ceiling, larger input limit, and broader voice library give it a narrow developer advantage.

Deepgram charges $30 per 1M characters (flagship) versus ElevenLabs at $100 per 1M characters, making it far cheaper at scale.

Deepgram charges $30 per 1M characters (flagship) versus ElevenLabs at $100 per 1M characters, making it far cheaper at scale. Deepgram also offers official SDKs in Python, JavaScript/TypeScript,.NET, and Go, compared to Python and JavaScript/TypeScript only from ElevenLabs. Its pure usage pricing model suits metered products well. ElevenLabs does edge ahead on word-level timestamps (yes versus no) and latency, with a 75 ms TTFB against Deepgram's 200 ms. Even so, on developer economics and SDK breadth, Deepgram wins.

For developers metering usage, MiniMax Speech offers word-level timestamps, pronunciation dictionaries, and 7 output formats including flac, opus, and pcmu variants.

For developers metering usage, MiniMax Speech offers word-level timestamps, pronunciation dictionaries, and 7 output formats including flac, opus, and pcmu variants. Fish Audio lacks verified timestamp support and charges $15 per 1M UTF-8 bytes, compared to MiniMax at $60-100 per 1M characters. MiniMax hybrid pricing also adds a $5 per month platform fee. Fish Audio wins on TTFB (100ms vs 250ms) and self-hosting. Overall, MiniMax's timestamps and richer output formats tip the developer tooling comparison narrowly in its favor.

Both APIs offer Python and Node/TypeScript SDKs, word-level timestamps, SSML, and streaming.
See pricingTry Murf API

Both APIs offer Python and Node/TypeScript SDKs, word-level timestamps, SSML, and streaming. Murf edges ahead on output flexibility with 5 formats (MP3, WAV, FLAC, ALAW, ULAW) versus Speechify's 2 (MP3, WAV), and supports a larger max input of 3000 characters versus 2000. Murf's pricing is pure usage-based, which suits metering, while Speechify's hybrid model requires a platform fee. Murf also offers a real-time WebSocket API and 5 concurrent requests on its base plan, though Speechify's Starter allows 15. The broader output format support and pure usage pricing give Murf a narrow developer advantage.

For developer integration, ElevenLabs offers word-level timestamps, SSML support, pronunciation dictionaries, and a real-time WebSocket API, none of which Mistral Voxtral TTS supports.

For developer integration, ElevenLabs offers word-level timestamps, SSML support, pronunciation dictionaries, and a real-time WebSocket API, none of which Mistral Voxtral TTS supports. ElevenLabs also covers 32 languages versus 9 for Mistral. Mistral wins on price at 16 dollars per 1M chars versus ElevenLabs at 100 dollars per 1M chars, and offers self-hosting. However, the breadth of developer-facing features in ElevenLabs, specifically timestamps, SSML, and WebSocket streaming, directly addresses the stated use case of metering and building speech into products.

ElevenLabs offers word-level timestamps, a 3,000-voice library versus OpenAI TTS's 13 built-in voices, instant and professional voice cloning, and a 40,000-character input limit versus 4,096 for OpenAI TTS. Both provide Python and JS/TS SDKs and support streaming and websocket APIs. On cost, OpenAI TTS charges $30 per 1M characters for its flagship tier, while ElevenLabs charges $100 per 1M characters, a real cost disadvantage for ElevenLabs. However, timestamps and cloning are strong developer differentiators that tip the balance narrowly toward ElevenLabs for feature-rich product builds.

For developers metering usage into products, ElevenLabs offers word-level timestamps, SSML support, pronunciation dictionaries, and a 40,000-character max input per request, features Fish Audio lacks in the verified facts. Both provide Python and TypeScript SDKs and streaming. Fish Audio's pure usage pricing at 15 $/1M bytes is cheaper than ElevenLabs at 100 $/1M chars, but ElevenLabs' tooling depth across timestamps, SSML, and pronunciation dictionaries gives developers more precise control over speech output, which is critical when building polished products.

ElevenLabs offers Python and JavaScript/TypeScript SDKs, word-level timestamps, SSML support, pronunciation dictionaries, a realtime websocket API, and 32 languages, all verified facts. Its usage-based pricing starts at $50/1M chars (fast model) with a metered hybrid model, making it straightforward to meter in a product. Dia/Dia2 has only a PyTorch pip package, is flagged as dormant, and supports only 1 language, making it unsuitable for production developer use.

Both tools offer Python and JS/TS SDKs, websocket APIs, word-level timestamps, and pronunciation dictionaries. ElevenLabs edges ahead on developer flexibility: it supports a broader output format range, including MP3 at multiple sample rates, versus Cartesia's raw PCM focus, handles up to 40,000 characters per request, and offers a hybrid pricing model that scales cleanly. ElevenLabs also has a 3,000-voice library giving developers more variety to meter and resell. Cartesia's 90 ms TTFB versus ElevenLabs' 75 ms slightly favors ElevenLabs for latency-sensitive products. The margins are close, but ElevenLabs wins on breadth of format options and request size.

For metered usage pricing, Rime charges a flat $50 per 1M characters with no platform fee, while Inworld TTS requires a $25 per month platform fee on top of usage costs.
See pricingTry Rime

For metered usage pricing, Rime charges a flat $50 per 1M characters with no platform fee, while Inworld TTS requires a $25 per month platform fee on top of usage costs. Rime offers 20 concurrent requests on its base plan versus Inworld's 5, making it better for scaling products. Both support word-level timestamps and streaming. Inworld has a Python SDK advantage, but Rime's higher concurrency, pure usage pricing model, and 120ms TTFB versus 200ms tip it narrowly toward developers building metered, latency-sensitive products.

For developers building speech products, output format flexibility and API feature depth matter.

For developers building speech products, output format flexibility and API feature depth matter. OpenAI TTS supports 6 output formats (MP3, Opus, AAC, FLAC, WAV, PCM) versus Mistral's 2 (PCM, MP3), offers a realtime WebSocket API, emotion and style controls, and a 4096-character input limit per request. Both have usage-based pricing and Python plus TypeScript SDKs. Mistral wins on flagship price ($16 vs. $30 per 1M characters) and instant voice cloning, but OpenAI's broader format support and WebSocket API are decisive developer advantages.

Cloud-utility TTS at commodity prices.

Cloud-utility TTS at commodity prices. No won verdicts for this use case yet; it ranks on ties and near-misses.

Best-known open model for scripted two-speaker dialogue rather than narration - Apache-2.0, English-only, GPU-oriented.
See pricingWebsite →

Best-known open model for scripted two-speaker dialogue rather than narration - Apache-2.0, English-only, GPU-oriented. No won verdicts for this use case yet; it ranks on ties and near-misses.

Developer-first hosted API from the team behind the open-source fish-speech models; simple prepaid pay-as-you-go billing, 80+ language coverage, and a free fair-use model tier (s2.1-pro-free). Facts here cover the hosted API, not the OSS model.

Developer-first hosted API from the team behind the open-source fish-speech models; simple prepaid pay-as-you-go billing, 80+ language coverage, and a free fair-use model tier (s2.1-pro-free). Facts here cover the hosted API, not the OSS model. No won verdicts for this use case yet; it ranks on ties and near-misses.

Cost-leader realtime TTS for voice agents and games; hybrid pay-as-you-go plus monthly credit plans that lower the per-1M-character rate as commitment grows; SOC 2 Type II with zero-data-retention posture.

Cost-leader realtime TTS for voice agents and games; hybrid pay-as-you-go plus monthly credit plans that lower the per-1M-character rate as commitment grows; SOC 2 Type II with zero-data-retention posture. No won verdicts for this use case yet; it ranks on ties and near-misses.

Speed-and-affordability challenger: a single fast model (Blizzard 2) with streaming-first APIs, unlimited voice clones on every tier, and flat subscription pricing with per-1K overage.
From $10/moTry LMNT

Speed-and-affordability challenger: a single fast model (Blizzard 2) with streaming-first APIs, unlimited voice clones on every tier, and flat subscription pricing with per-1K overage. No won verdicts for this use case yet; it ranks on ties and near-misses.

Developer platform spun out of the Speechify brand: transparent tiered pricing ($10-$499/mo plus per-1M overage), streaming-native Simba 3.2, and a bundled voice-agents product with flat per-minute rates.

Developer platform spun out of the Speechify brand: transparent tiered pricing ($10-$499/mo plus per-1M overage), streaming-native Simba 3.2, and a bundled voice-agents product with flat per-minute rates. No won verdicts for this use case yet; it ranks on ties and near-misses.

Open-weights-friendly voice cloning TTS from a frontier AI lab.

Open-weights-friendly voice cloning TTS from a frontier AI lab. No won verdicts for this use case yet; it ranks on ties and near-misses.

More Text-to-speech APIs buyer guides