# Best Text-to-speech APIs for Self-Hosted (2026)

> The best Text-to-speech APIs platforms for self-hosted: Voxtral TTS leads, for running an open-weight model on your own hardware: license reality, model size.

For self-hosted, **Voxtral TTS** is our pick: Mistral Voxtral TTS offers a verified self-host option and releases model weights under CC BY-NC 4.0, allowing users to run it on their own hardware. Running an open-weight model on your own hardware: license reality, model size, hardware needs, and whether the project is still alive. Below is the full ranking and the tradeoffs, or read [how we score](https://www.versusref.com/methodology/).

## What matters for self-hosted

Weight ×5 = decisive, ×1 = relevant.

| Fact | Weight | Voxtral TTS | CAMB.AI | CosyVoice | Fish Audio | Rime |
| --- | --- | --- | --- | --- | --- | --- |
| Self-host / on-prem option | ×5 | ✓  Yes (Jul 20) | n/a | ✓  Yes (Jul 20) | ✓  Yes (Jul 20) | ✓  Yes (Jul 20) |
| Model weights license | ×5 | CC BY-NC 4.0 (open weights, non-commercial) (Jul 20) | Open-source (license unstated) (Jul 20) | Apache-2.0 (Jul 20) | n/a | n/a |
| Hardware to self-host | ×4 | n/a | n/a | Python 3.10 (Jul 20) | n/a | n/a |
| Project maintenance status | ×3 | n/a | n/a | active (Jul 20) | n/a | n/a |
| Model size (parameters) | ×2 | n/a | n/a | 0.5B (CosyVoice 2.0 and Fun-CosyVoice 3.0) (Jul 20) | n/a | n/a |

- ×5 **Self-host / on-prem option:** This page only ranks what you can actually run yourself.
- ×5 **Model weights license:** The license decides commercial viability: Apache/MIT ship products, non-commercial licenses do not.
- ×4 **Hardware to self-host:** CPU-capable small models and 16GB-VRAM giants are different budgets entirely.
- ×3 **Project maintenance status:** A dormant repo means you own every future bug and compatibility break.
- ×2 **Model size (parameters):** Parameter count is the quickest proxy for speed/quality trade-off on your hardware.

## The ranking, tool by tool

| Rank | Tool | Verdict | Score | Price |
| --- | --- | --- | --- | --- |
| 1 | [Voxtral TTS](https://www.versusref.com/tts/tools/voxtral-tts/) | Mistral Voxtral TTS offers a verified self-host option and releases model weights under CC BY-NC 4.0, allowing users to run it on their own hardware. | 6 of 6 points · 3 matchups | See pricing |
| 2 | [CAMB.AI](https://www.versusref.com/tts/tools/camb-ai/) | For self-hosting, the critical factor is whether model weights are available. | 2 of 2 points · 1 matchup | From $5/mo |
| 3 | [CosyVoice](https://www.versusref.com/tts/tools/cosyvoice/) (OSS) | For self-hosting, licensing is critical. | 2 of 2 points · 1 matchup | See pricing |
| 4 | [Fish Audio](https://www.versusref.com/tts/tools/fish-audio/) | Fish Audio explicitly supports a self-host/on-prem option, while MiniMax Speech offers no self-host option at all. | 2 of 2 points · 1 matchup | See pricing |
| 5 | [Rime](https://www.versusref.com/tts/tools/rime/) | Rime has a verified self-host and on-premises option, while no such capability is recorded for Inworld TTS. | 2.5 of 4 points · 2 matchups | See pricing |
| 6 | [Azure Speech](https://www.versusref.com/tts/tools/azure-speech/) | For a self-hosted deployment, the two most critical factors are whether the vendor offers an on-premises option and what license governs the model weights. | 4.5 of 8 points · 4 matchups | From $960/mo |
| 7 | [Dia / Dia2](https://www.versusref.com/tts/tools/dia/) (OSS) | Dia2 ships under Apache-2.0 with open weights (7bd5a6ab), supports self-hosting on GPU hardware (ce3d9d0b), and offers 1B and 2B parameter model sizes suitable for local deployment (d4c6fd12). | 1 of 2 points · 1 matchup | See pricing |
| 8 | [Deepgram Aura-2](https://www.versusref.com/tts/tools/deepgram-aura/) | For self-hosted deployment, Deepgram Aura-2 TTS explicitly supports a self-host or on-prem option per fact 039d1788, giving it a concrete deployment path that ElevenLabs cannot match in the verified facts. | 1.5 of 4 points · 2 matchups | See pricing |
| 9 | [Cartesia](https://www.versusref.com/tts/tools/cartesia/) | The use case asks about self-hosting capability. | 5 of 14 points · 7 matchups | From $5/mo |
| 10 | [Amazon Polly](https://www.versusref.com/tts/tools/amazon-polly/) | Cloud-utility TTS at commodity prices. | 1 of 4 points · 2 matchups | See pricing |
| 11 | [Speechify API](https://www.versusref.com/tts/tools/speechify-api/) | Developer platform spun out of the Speechify brand: transparent tiered pricing ($10-$499/mo plus per-1M overage), streaming-native Simba 3.2, and a bundled voice-agents product with flat per-minute rates. | 0.5 of 2 points · 1 matchup | From $10/mo |
| 12 | [Google Cloud TTS](https://www.versusref.com/tts/tools/google-tts/) | Hyperscaler TTS with the broadest voice/language catalog. | 1.5 of 8 points · 4 matchups | See pricing |
| 13 | [Murf API](https://www.versusref.com/tts/tools/murf/) | API arm of the Murf studio platform: 150+ voices in 35 languages, SSML support, word timestamps, and Falcon 2 aimed at high-concurrency voice agents at $0.01/1K characters. Note: Murf Studio subscription plans (murf.ai/pricing) are a separate product from API pay-as-you-go pricing and API characters are purchased separately. | 1 of 6 points · 3 matchups | See pricing |
| 14 | [ElevenLabs](https://www.versusref.com/tts/tools/elevenlabs/) | Premium AI voice platform for creators and developers. | 2.5 of 20 points · 10 matchups | From $6/mo |
| 15 | [OpenAI TTS](https://www.versusref.com/tts/tools/openai-tts/) | Simple usage-based TTS inside a general AI platform. | 1 of 10 points · 5 matchups | See pricing |
| 16 | [Fish Speech](https://www.versusref.com/tts/tools/fish-speech/) (OSS) | Top-tier expressive multilingual open-weights TTS whose license moved from permissive to research/non-commercial; commercial use requires a license from Fish Audio or their hosted API. | 0 of 2 points · 1 matchup | See pricing |
| 17 | [Inworld TTS](https://www.versusref.com/tts/tools/inworld-tts/) | Cost-leader realtime TTS for voice agents and games; hybrid pay-as-you-go plus monthly credit plans that lower the per-1M-character rate as commitment grows; SOC 2 Type II with zero-data-retention posture. | 0 of 4 points · 2 matchups | From $25/mo |
| 18 | [LMNT](https://www.versusref.com/tts/tools/lmnt/) | Speed-and-affordability challenger: a single fast model (Blizzard 2) with streaming-first APIs, unlimited voice clones on every tier, and flat subscription pricing with per-1K overage. | 0 of 2 points · 1 matchup | From $10/mo |
| 19 | [MiniMax Speech](https://www.versusref.com/tts/tools/minimax-speech/) | Multilingual cloning-first TTS with aggressive pricing. | 0 of 2 points · 1 matchup | From $5/mo |

### 1. Voxtral TTS

Mistral Voxtral TTS offers a verified self-host option and releases model weights under CC BY-NC 4.0, allowing users to run it on their own hardware. [Full Voxtral TTS vs OpenAI TTS verdict](https://www.versusref.com/tts/openai-tts-vs-voxtral-tts/)

For self-hosted deployment, Mistral Voxtral TTS wins on every relevant dimension. [Full Voxtral TTS vs ElevenLabs verdict](https://www.versusref.com/tts/elevenlabs-vs-voxtral-tts/)

For self-hosted open-weight deployment, Mistral Voxtral TTS ships with CC BY-NC 4.0 open weights, meaning the model weights are publicly available for download and self-hosting on your own hardware. [Full Voxtral TTS vs Cartesia verdict](https://www.versusref.com/tts/cartesia-vs-voxtral-tts/)

### 2. CAMB.AI

For self-hosting, the critical factor is whether model weights are available. [Full CAMB.AI vs ElevenLabs verdict](https://www.versusref.com/tts/camb-ai-vs-elevenlabs/)

### 3. CosyVoice

For self-hosting, licensing is critical. [Full CosyVoice vs Fish Speech verdict](https://www.versusref.com/tts/cosyvoice-vs-fish-speech/)

### 4. Fish Audio

Fish Audio explicitly supports a self-host/on-prem option, while MiniMax Speech offers no self-host option at all. [Full Fish Audio vs MiniMax Speech verdict](https://www.versusref.com/tts/fish-audio-vs-minimax-speech/)

### 5. Rime

Rime has a verified self-host and on-premises option, while no such capability is recorded for Inworld TTS. [Full Rime vs Inworld TTS verdict](https://www.versusref.com/tts/inworld-tts-vs-rime/)

### 6. Azure Speech

For a self-hosted deployment, the two most critical factors are whether the vendor offers an on-premises option and what license governs the model weights. [Full Azure Speech vs Murf API verdict](https://www.versusref.com/tts/azure-speech-vs-murf/)

The use case asks about self-hosting on own hardware. [Full Azure Speech vs OpenAI TTS verdict](https://www.versusref.com/tts/azure-speech-vs-openai-tts/)

Both tools use closed model weights, so neither supports true open-weight self-hosting. [Full Azure Speech vs Google Cloud TTS verdict](https://www.versusref.com/tts/azure-speech-vs-google-tts/)

### 7. Dia / Dia2

Dia2 ships under Apache-2.0 with open weights (7bd5a6ab), supports self-hosting on GPU hardware (ce3d9d0b), and offers 1B and 2B parameter model sizes suitable for local deployment (d4c6fd12). [Full Dia / Dia2 vs ElevenLabs verdict](https://www.versusref.com/tts/dia-vs-elevenlabs/)

### 8. Deepgram Aura-2

For self-hosted deployment, Deepgram Aura-2 TTS explicitly supports a self-host or on-prem option per fact 039d1788, giving it a concrete deployment path that ElevenLabs cannot match in the verified facts. [Full Deepgram Aura-2 vs ElevenLabs verdict](https://www.versusref.com/tts/deepgram-aura-vs-elevenlabs/)

### 9. Cartesia

The use case asks about self-hosting capability. [Full Cartesia vs OpenAI TTS verdict](https://www.versusref.com/tts/cartesia-vs-openai-tts/)

Cartesia (Sonic) explicitly supports self-hosting or on-premises deployment per fact 57210ab8, making it directly relevant to this use case. [Full Cartesia vs LMNT verdict](https://www.versusref.com/tts/cartesia-vs-lmnt/)

Cartesia Sonic explicitly supports a self-host and on-prem option per fact 57210ab8, giving it a concrete deployment path for self-hosted scenarios. [Full Cartesia vs Inworld TTS verdict](https://www.versusref.com/tts/cartesia-vs-inworld-tts/)

Both tools have closed model weight licenses, so neither supports true open-weight self-hosting. [Full Cartesia vs ElevenLabs verdict](https://www.versusref.com/tts/cartesia-vs-elevenlabs/)

### 10. Amazon Polly

Cloud-utility TTS at commodity prices. No won verdicts for this use case yet; it ranks on ties and near-misses.

### 11. Speechify API

Developer platform spun out of the Speechify brand: transparent tiered pricing ($10-$499/mo plus per-1M overage), streaming-native Simba 3.2, and a bundled voice-agents product with flat per-minute rates. No won verdicts for this use case yet; it ranks on ties and near-misses.

### 12. Google Cloud TTS

Hyperscaler TTS with the broadest voice/language catalog. No won verdicts for this use case yet; it ranks on ties and near-misses.

### 13. Murf API

API arm of the Murf studio platform: 150+ voices in 35 languages, SSML support, word timestamps, and Falcon 2 aimed at high-concurrency voice agents at $0.01/1K characters. Note: Murf Studio subscription plans (murf.ai/pricing) are a separate product from API pay-as-you-go pricing and API characters are purchased separately. No won verdicts for this use case yet; it ranks on ties and near-misses.

### 14. ElevenLabs

Premium AI voice platform for creators and developers. No won verdicts for this use case yet; it ranks on ties and near-misses.

### 15. OpenAI TTS

Simple usage-based TTS inside a general AI platform. No won verdicts for this use case yet; it ranks on ties and near-misses.

### 16. Fish Speech

Top-tier expressive multilingual open-weights TTS whose license moved from permissive to research/non-commercial; commercial use requires a license from Fish Audio or their hosted API. No won verdicts for this use case yet; it ranks on ties and near-misses.

### 17. Inworld TTS

Cost-leader realtime TTS for voice agents and games; hybrid pay-as-you-go plus monthly credit plans that lower the per-1M-character rate as commitment grows; SOC 2 Type II with zero-data-retention posture. No won verdicts for this use case yet; it ranks on ties and near-misses.

### 18. LMNT

Speed-and-affordability challenger: a single fast model (Blizzard 2) with streaming-first APIs, unlimited voice clones on every tier, and flat subscription pricing with per-1K overage. No won verdicts for this use case yet; it ranks on ties and near-misses.

### 19. MiniMax Speech

Multilingual cloning-first TTS with aggressive pricing. No won verdicts for this use case yet; it ranks on ties and near-misses.

Source: https://www.versusref.com/tts/best/self-hosted/
