vsref
XTTS v2 (Coqui) logo

XTTS v2 (Coqui) Review

OSS

The legacy standard for open voice cloning: dormant upstream since Coqui's Jan 2024 shutdown (community fork idiap/coqui-ai-TTS carries maintenance), and CPML weights bar commercial use.

Among the 46 text-to-speech tools we track, XTTS v2 (Coqui) has the 21st-widest language coverage.

See pricing

Facts verified Jul 20, 2026Website →

What we know about XTTS v2 (Coqui)

XTTS v2 (Coqui) is a text-to-speech apis platform: the legacy standard for open voice cloning: dormant upstream since Coqui's Jan 2024 shutdown (community fork idiap/coqui-ai-TTS carries maintenance), and CPML weights bar commercial use. This profile tracks every XTTS v2 (Coqui) fact we have verified, each linked to a primary source and dated.

XTTS v2 (Coqui) does not publish a public per-minute rate we have been able to verify, so the price figure here stays blank until we can confirm one. We would rather show nothing than a guessed number; if you have a current XTTS v2 (Coqui) quote, it is the fastest way for us to close that gap.

On capabilities, XTTS v2 (Coqui) covers instant voice cloning, emotion / style controls, and self-host / on-prem option. Each of those is verified against XTTS v2 (Coqui)'s own docs or dashboard, not marketing copy.

Placed against the 46 text-to-speech tools we track, XTTS v2 (Coqui)'s strongest showing is the 21st-widest language coverage - a spread worth weighing against your own priorities.

In total we track 10 verified facts for XTTS v2 (Coqui) today, each linking the primary source it came from so you can check our work - and vendor claims we have not measured ourselves are labeled as such on the XTTS v2 (Coqui) fact sheet below.

Reviewed by vsref Editorialfacts verified Jul 20, 2026Methodology →

Fact sheet

Capabilities
Capabilities facts
Instant voice cloning✓ YesJul 20
Languages supported17 languagesJul 20
Emotion / style controls✓ YesJul 20
Compliance & trust
Compliance & trust facts
Self-host / on-prem option✓ YesJul 20
Model weights licenseCPML (Coqui Public Model License)Jul 20
Build experience
Build experience facts
Official SDKsCoqui TTS Python library (pip `TTS`; fork: pip `coqui-tts`)Jul 20
Output formats24 kHz audio outputJul 20
Commercial
Commercial facts
Hosted API availableNoJul 20
Project maintenance statusdormantJul 20
GitHub stars45,800 starsJul 20

Considering a switch? Best XTTS v2 (Coqui) alternatives →

STT in this stack

Text-to-speech is half of a voice pipeline: the other half is the speech-to-text that listens. Compare transcription engines on accuracy, streaming latency, and per-minute price: Deepgram, AssemblyAI, GPT-4o Transcribe, or the full speech-to-text comparison →

Voice agents in this stack

The engine is one layer: a voice agent speaks through its text-to-speech engine, but orchestration, telephony, and turn-taking come from the agent platform. Compare the platforms builders pair XTTS v2 (Coqui) with: Pipecat, OpenAI Realtime API, Twilio ConversationRelay, or the full voice-agent comparison →

Distribute it

Most XTTS v2 (Coqui) voiceover ends up in short-form video, and publishing that video across TikTok, YouTube, and Instagram is a scheduling problem with real per-channel pricing. Compare the schedulers creators actually run: Buffer, Postiz, Mixpost, or the social media scheduling platforms compared →

Avatar video in this stack

A cloned or bring-your-own XTTS v2 (Coqui) voice does not have to stay audio-only: AI avatar video platforms lip-sync it onto a talking avatar for finished video. Compare the platforms: HeyGen, Synthesia, Hedra, or the full avatar-video comparison →