vsref
CosyVoice logo

CosyVoice

OSS

Full-stack open TTS from Alibaba's speech team: 9 languages plus 18+ Chinese dialects, 150 ms streaming latency claim, instruction control of emotion/dialect/speed, and training + deployment scripts under Apache-2.0.

Among the 46 text-to-speech tools we track, CosyVoice has the 26th-widest language coverage.

See pricing

Facts verified Jul 20, 2026Website →

What we know about CosyVoice

CosyVoice is a text-to-speech apis platform: full-stack open TTS from Alibaba's speech team: 9 languages plus 18+ Chinese dialects, 150 ms streaming latency claim, instruction control of emotion/dialect/speed, and training + deployment scripts under Apache-2.0. This profile tracks what we have verified about CosyVoice, each fact linked to a primary source and dated.

CosyVoice does not publish a public per-minute rate we have been able to verify, so the price figure here stays blank until we can confirm one. We would rather show nothing than a guessed number; if you have a current CosyVoice quote, it is the fastest way for us to close that gap.

On capabilities, CosyVoice covers streaming audio output, instant voice cloning, emotion / style controls, and self-host / on-prem option. Each of those is verified against CosyVoice's own docs or dashboard, not marketing copy.

In total we track 11 verified facts for CosyVoice today, and add coverage as the ingestion pass revisits it. Where a number is a vendor claim rather than our own measurement, it is labeled as such on the CosyVoice fact sheet below, and each row links the primary source it came from so you can check our work.

Whether CosyVoice is the right call depends on your use case more than any single spec, which is why the head-to-head verdicts below score it per scenario rather than crowning one overall winner. Use the fact sheet for the raw numbers, and the matchups for how CosyVoice actually fares against the platforms buyers most often weigh it against.

Fact sheet

Capabilities
Capabilities facts
Streaming audio output✓ YesJul 20
Instant voice cloning✓ YesJul 20
Languages supported9 languagesJul 20
Emotion / style controls✓ YesJul 20
Compliance & trust
Compliance & trust facts
Self-host / on-prem option✓ YesJul 20
Model weights licenseApache-2.0Jul 20
Build experience
Build experience facts
Official SDKsPython; gRPC/FastAPI deployment examples; DockerJul 20
Commercial
Commercial facts
Model size (parameters)0.5B (CosyVoice 2.0 and Fun-CosyVoice 3.0)Jul 20
Hardware to self-hostPython 3.10Jul 20
Project maintenance statusactiveJul 20
GitHub stars22,288 starsJul 20

Considering a switch? Best CosyVoice alternatives →

CosyVoice head-to-head

CosyVoice vs Fish Speechwon 1 · lost 3 · tied 0
AudiobookslostDubbinglostSelf-HostedwonVoice Agentslost