XTTS v2 (Coqui)
OSSThe legacy standard for open voice cloning: dormant upstream since Coqui's Jan 2024 shutdown (community fork idiap/coqui-ai-TTS carries maintenance), and CPML weights bar commercial use.
Among the 46 text-to-speech tools we track, XTTS v2 (Coqui) has the 21st-widest language coverage.
See pricing
What we know about XTTS v2 (Coqui)
XTTS v2 (Coqui) is a text-to-speech apis platform: the legacy standard for open voice cloning: dormant upstream since Coqui's Jan 2024 shutdown (community fork idiap/coqui-ai-TTS carries maintenance), and CPML weights bar commercial use. This profile tracks what we have verified about XTTS v2 (Coqui), each fact linked to a primary source and dated.
XTTS v2 (Coqui) does not publish a public per-minute rate we have been able to verify, so the price figure here stays blank until we can confirm one. We would rather show nothing than a guessed number; if you have a current XTTS v2 (Coqui) quote, it is the fastest way for us to close that gap.
On capabilities, XTTS v2 (Coqui) covers instant voice cloning, emotion / style controls, and self-host / on-prem option. Each of those is verified against XTTS v2 (Coqui)'s own docs or dashboard, not marketing copy.
In total we track 10 verified facts for XTTS v2 (Coqui) today, and add coverage as the ingestion pass revisits it. Where a number is a vendor claim rather than our own measurement, it is labeled as such on the XTTS v2 (Coqui) fact sheet below, and each row links the primary source it came from so you can check our work.
Whether XTTS v2 (Coqui) is the right call depends on your use case more than any single spec, which is why the head-to-head verdicts below score it per scenario rather than crowning one overall winner. Use the fact sheet for the raw numbers, and the matchups for how XTTS v2 (Coqui) actually fares against the platforms buyers most often weigh it against.
Fact sheet
Every row independently verifiedConsidering a switch? Best XTTS v2 (Coqui) alternatives →