vsref

Best Text-to-speech APIs for Self-Hosted (2026)

For self-hosted, Voxtral TTS is our pick: Mistral Voxtral TTS offers a verified self-host option and releases model weights under CC BY-NC 4.0, allowing users to run it on their own hardware. Running an open-weight model on your own hardware: license reality, model size, hardware needs, and whether the project is still alive. Below is the full ranking and the tradeoffs, or read how we score.

If you sign up through links on this page, vsref may earn a commission; programs exist on both sides of most comparisons, and commissions never change verdicts. How we make money

Open-weights-friendly voice cloning TTS from a frontier AI lab6 of 6 points · 3 matchups
Localization-first TTS: the MARS 8 family (flash/pro/instruct variants) plus dubbing and translated-TTS pipelines, credit-based plans from $5/mo, aimed at media, sports, and content going multilingual.2 of 2 points · 1 matchup
From $5/moTry CAMB.AI
Full-stack open TTS from Alibaba's speech team: 9 languages plus 18+ Chinese dialects, 150 ms streaming latency claim, instruction control of emotion/dialect/speed, and training + deployment scripts under Apache-2.0.2 of 2 points · 1 matchup
See pricingWebsite →

What matters for self-hosted

Weighted attribute comparison for Self-Hosted
FactVoxtral TTSCAMB.AICosyVoiceFish AudioRime
Self-host / on-prem option×5✓ YesJul 20n/a✓ YesJul 20✓ YesJul 20✓ YesJul 20
Model weights license×5CC BY-NC 4.0 (open weights, non-commercial)Jul 20Open-source (license unstated)Jul 20Apache-2.0Jul 20n/an/a
Hardware to self-host×4n/an/aPython 3.10Jul 20n/an/a
Project maintenance status×3n/an/aactiveJul 20n/an/a
Model size (parameters)×2n/an/a0.5B (CosyVoice 2.0 and Fun-CosyVoice 3.0)Jul 20n/an/a
Swipe → to see every tool column.
×5 Self-host / on-prem option: This page only ranks what you can actually run yourself.×5 Model weights license: The license decides commercial viability: Apache/MIT ship products, non-commercial licenses do not.×4 Hardware to self-host: CPU-capable small models and 16GB-VRAM giants are different budgets entirely.×3 Project maintenance status: A dormant repo means you own every future bug and compatibility break.×2 Model size (parameters): Parameter count is the quickest proxy for speed/quality trade-off on your hardware.

The ranking, tool by tool

Mistral Voxtral TTS offers a verified self-host option and releases model weights under CC BY-NC 4.0, allowing users to run it on their own hardware.

Mistral Voxtral TTS offers a verified self-host option and releases model weights under CC BY-NC 4.0, allowing users to run it on their own hardware. OpenAI TTS has closed model weights and no self-host option. For the self-hosted use case, the facts point decisively to Mistral Voxtral TTS. The CC BY-NC license does restrict commercial use, but it remains a real, deployable open-weight model, compared to OpenAI's fully closed offering.

For self-hosted deployment, Mistral Voxtral TTS wins on every relevant dimension. It explicitly supports a self-host and on-premises option, and its weights are released under CC BY-NC 4.0, meaning users can run it on their own hardware legally. ElevenLabs has closed model weights and no self-host option. The CC BY-NC 4.0 license does restrict commercial use, but for on-premises deployment this is a decisive advantage over a fully closed model that cannot be self-hosted at all.

For self-hosted open-weight deployment, Mistral Voxtral TTS ships with CC BY-NC 4.0 open weights, meaning the model weights are publicly available for download and self-hosting on your own hardware. Cartesia Sonic has closed model weights, so true self-hosting of the model itself is not possible, even though Cartesia offers an on-prem API option. When the use case explicitly requires running an open-weight model on your own hardware, Mistral's open weights license wins decisively over Cartesia's closed weights.

For self-hosting, the critical factor is whether model weights are available.
From $5/moTry CAMB.AI

For self-hosting, the critical factor is whether model weights are available. CAMB.AI offers open-source model weights, while ElevenLabs uses a fully closed model with no self-hosting option. CAMB.AI also has an official Go SDK alongside Python and Node.js, adding deployment flexibility. ElevenLabs provides no open weights, making self-hosted deployment impossible regardless of hardware or budget.

For self-hosting, licensing is critical.
See pricingWebsite →

For self-hosting, licensing is critical. CosyVoice uses Apache-2.0, permitting commercial use freely, while Fish Speech uses a custom non-commercial research license, blocking most production deployments. CosyVoice also runs on a lighter 0.5B model with Python 3.10 as the main hardware requirement, versus Fish Speech requiring a GPU with benchmarks on an NVIDIA H200 and a 4B model. Both are actively maintained, but the permissive license and lower hardware bar make CosyVoice the clear winner for practical self-hosting.

Fish Audio explicitly supports a self-host/on-prem option, while MiniMax Speech offers no self-host option at all.

Fish Audio explicitly supports a self-host/on-prem option, while MiniMax Speech offers no self-host option at all. MiniMax also uses closed model weights, making any on-prem deployment impossible regardless. Fish Audio wins decisively on both the deployment flexibility and model openness dimensions that define this use case.

Rime has a verified self-host and on-premises option, while no such capability is recorded for Inworld TTS.
See pricingTry Rime

Rime has a verified self-host and on-premises option, while no such capability is recorded for Inworld TTS. For a use case centered on running software on your own hardware, this is the decisive differentiator. Because no comparable self-hosting fact exists for Inworld TTS in the provided fact set, Rime is the clear winner.

Dia2 ships under Apache-2.0 with open weights (7bd5a6ab), supports self-hosting on GPU hardware (ce3d9d0b), and offers 1B and 2B parameter model sizes suitable for local deployment (d4c6fd12).
See pricingWebsite →

Dia2 ships under Apache-2.0 with open weights, supports self-hosting on GPU hardware, and offers 1B and 2B parameter model sizes suitable for local deployment. ElevenLabs has closed weights, making self-hosting impossible. The one meaningful concern is that Dia2's project maintenance is marked dormant, which is a real risk. However, since ElevenLabs cannot be self-hosted at all, Dia2 wins by default despite the dormancy caveat.

The use case asks about self-hosting on own hardware.

The use case asks about self-hosting on own hardware. Azure Speech explicitly supports a self-host or on-premises deployment option, while the available facts mention no equivalent capability for OpenAI TTS. That said, both tools use closed model weights, so neither offers true open-weight self-hosting. Azure still wins narrowly because it at least has a documented self-host deployment path, giving it a practical edge over OpenAI TTS, which has no self-host option mentioned at all.

Both tools use closed model weights, so neither supports true open-weight self-hosting. However, Azure Speech explicitly offers a self-host or on-prem deployment option, while no equivalent fact exists for Google Cloud TTS. Google TTS model weights are confirmed closed. Azure's closed weights are also confirmed, but its on-prem option still gives it a concrete advantage for teams needing local deployment under data-residency or air-gap requirements.

For self-hosted deployment, Deepgram Aura-2 TTS explicitly supports a self-host or on-prem option per fact 039d1788, giving it a concrete deployment path that ElevenLabs cannot match in the verified facts.

For self-hosted deployment, Deepgram Aura-2 TTS explicitly supports a self-host or on-prem option, giving it a concrete deployment path that ElevenLabs cannot match in the verified facts. Both tools use closed model weights, so neither allows true open-weight local inference. Even so, Deepgram wins narrowly because having an official self-host option beats having no verified self-host path at all, despite both sharing the closed-weights limitation.

The use case asks about self-hosting capability.

The use case asks about self-hosting capability. Cartesia Sonic explicitly supports a self-host or on-prem option, while OpenAI TTS provides no such option in the verified facts. Both tools have closed model weights, so neither offers true open-weight self-hosting. Cartesia wins narrowly because at least a self-hosted deployment path exists, even if the weights are proprietary and the arrangement is likely enterprise.

Cartesia (Sonic) explicitly supports self-hosting or on-premises deployment, making it directly relevant to this use case. LMNT has no verified self-host fact in the provided data. Cartesia's model weights are closed-license, meaning users cannot freely inspect or redistribute them, which limits true open-weight self-hosting. Neither tool offers open weights, but Cartesia at least has a documented self-host path, giving it a narrow edge over LMNT, which has no self-host support recorded at all.

Cartesia Sonic explicitly supports a self-host and on-prem option, giving it a concrete deployment path for self-hosted scenarios. Both tools use closed model weights (the data confirms Cartesia is closed, and no open-weights fact exists for Inworld either), so neither allows running true open-weight models on personal hardware in the traditional sense. Cartesia at least has a documented self-host offering, while no such fact exists for Inworld, making Cartesia the narrowly better fit for self-hosted enterprise use despite its closed weights.

Both tools have closed model weight licenses, so neither supports true open-weight self-hosting. However, Cartesia Sonic explicitly offers a self-host or on-prem option, while no equivalent fact exists for ElevenLabs. Both have closed weights, meaning neither lets you run freely licensed weights, but Cartesia at least provides a formal self-hosted deployment path. The use case asks about running on your own hardware, and only Cartesia has a verified self-host option.

Cloud-utility TTS at commodity prices.

Cloud-utility TTS at commodity prices. No won verdicts for this use case yet; it ranks on ties and near-misses.

API arm of the Murf studio platform: 150+ voices in 35 languages, SSML support, word timestamps, and Falcon 2 aimed at high-concurrency voice agents at $0.01/1K characters. Note: Murf Studio subscription plans (murf.ai/pricing) are a separate product from API pay-as-you-go pricing and API characters are purchased separately.
See pricingTry Murf API

API arm of the Murf studio platform: 150+ voices in 35 languages, SSML support, word timestamps, and Falcon 2 aimed at high-concurrency voice agents at $0.01/1K characters. Note: Murf Studio subscription plans (murf.ai/pricing) are a separate product from API pay-as-you-go pricing and API characters are purchased separately. No won verdicts for this use case yet; it ranks on ties and near-misses.

Developer platform spun out of the Speechify brand: transparent tiered pricing ($10-$499/mo plus per-1M overage), streaming-native Simba 3.2, and a bundled voice-agents product with flat per-minute rates.

Developer platform spun out of the Speechify brand: transparent tiered pricing ($10-$499/mo plus per-1M overage), streaming-native Simba 3.2, and a bundled voice-agents product with flat per-minute rates. No won verdicts for this use case yet; it ranks on ties and near-misses.

Hyperscaler TTS with the broadest voice/language catalog.

Hyperscaler TTS with the broadest voice/language catalog. No won verdicts for this use case yet; it ranks on ties and near-misses.

Premium AI voice platform for creators and developers.

Premium AI voice platform for creators and developers. No won verdicts for this use case yet; it ranks on ties and near-misses.

Simple usage-based TTS inside a general AI platform.

Simple usage-based TTS inside a general AI platform. No won verdicts for this use case yet; it ranks on ties and near-misses.

Top-tier expressive multilingual open-weights TTS whose license moved from permissive to research/non-commercial; commercial use requires a license from Fish Audio or their hosted API.
See pricingWebsite →

Top-tier expressive multilingual open-weights TTS whose license moved from permissive to research/non-commercial; commercial use requires a license from Fish Audio or their hosted API. No won verdicts for this use case yet; it ranks on ties and near-misses.

Cost-leader realtime TTS for voice agents and games; hybrid pay-as-you-go plus monthly credit plans that lower the per-1M-character rate as commitment grows; SOC 2 Type II with zero-data-retention posture.

Cost-leader realtime TTS for voice agents and games; hybrid pay-as-you-go plus monthly credit plans that lower the per-1M-character rate as commitment grows; SOC 2 Type II with zero-data-retention posture. No won verdicts for this use case yet; it ranks on ties and near-misses.

Speed-and-affordability challenger: a single fast model (Blizzard 2) with streaming-first APIs, unlimited voice clones on every tier, and flat subscription pricing with per-1K overage.
From $10/moTry LMNT

Speed-and-affordability challenger: a single fast model (Blizzard 2) with streaming-first APIs, unlimited voice clones on every tier, and flat subscription pricing with per-1K overage. No won verdicts for this use case yet; it ranks on ties and near-misses.

Multilingual cloning-first TTS with aggressive pricing.

Multilingual cloning-first TTS with aggressive pricing. No won verdicts for this use case yet; it ranks on ties and near-misses.

More Text-to-speech APIs buyer guides