# CosyVoice Review

> CosyVoice review: Full-stack open TTS from Alibaba's speech team: 9 languages plus 18+ Chinese dialects, 150 ms streaming latency claim, instruction control of.

![CosyVoice logo](https://www.versusref.com/logos/cosyvoice.png)

Full-stack open TTS from Alibaba's speech team: 9 languages plus 18+ Chinese dialects, 150 ms streaming latency claim, instruction control of emotion/dialect/speed, and training + deployment scripts under Apache-2.0.

Among the 46 text-to-speech tools we track, CosyVoice has the 26th-widest language coverage.

See pricing · Open source · facts verified Jul 20, 2026

## What we know about CosyVoice

CosyVoice sits in the text-to-speech apis category, where it is full-stack open TTS from Alibaba's speech team: 9 languages plus 18+ Chinese dialects, 150 ms streaming latency claim, instruction control of emotion/dialect/speed, and training + deployment scripts under Apache-2.0. We keep this CosyVoice profile grounded in primary sources, each fact dated to when we last confirmed it.

CosyVoice does not publish a public per-minute rate we have been able to verify, so the price figure here stays blank until we can confirm one. We would rather show nothing than a guessed number; if you have a current CosyVoice quote, it is the fastest way for us to close that gap.

On capabilities, CosyVoice covers streaming audio output, instant voice cloning, emotion / style controls, and self-host / on-prem option. Each of those is verified against CosyVoice's own docs or dashboard, not marketing copy.

Placed against the 46 text-to-speech tools we track, CosyVoice's strongest showing is the 26th-widest language coverage - a spread worth weighing against your own priorities.

In total we track 11 verified facts for CosyVoice today, each linking the primary source it came from so you can check our work - and vendor claims we have not measured ourselves are labeled as such on the CosyVoice fact sheet below.

## Fact sheet

### Capabilities

| Fact | Value | Verified | Source |
| --- | --- | --- | --- |
| Streaming audio output | ✓  Yes | Jul 20 | [source](https://github.com/FunAudioLLM/CosyVoice) |
| Instant voice cloning | ✓  Yes | Jul 20 | [source](https://github.com/FunAudioLLM/CosyVoice) |
| Languages supported | 9 languages | Jul 20 | [source](https://github.com/FunAudioLLM/CosyVoice) |
| Emotion / style controls | ✓  Yes | Jul 20 | [source](https://github.com/FunAudioLLM/CosyVoice) |

### Compliance & trust

| Fact | Value | Verified | Source |
| --- | --- | --- | --- |
| Self-host / on-prem option | ✓  Yes | Jul 20 | [source](https://github.com/FunAudioLLM/CosyVoice) |
| Model weights license | Apache-2.0 | Jul 20 | [source](https://github.com/FunAudioLLM/CosyVoice) |

### Build experience

| Fact | Value | Verified | Source |
| --- | --- | --- | --- |
| Official SDKs | Python; gRPC/FastAPI deployment examples; Docker | Jul 20 | [source](https://github.com/FunAudioLLM/CosyVoice) |

### Commercial

| Fact | Value | Verified | Source |
| --- | --- | --- | --- |
| Model size (parameters) | 0.5B (CosyVoice 2.0 and Fun-CosyVoice 3.0) | Jul 20 | [source](https://github.com/FunAudioLLM/CosyVoice) |
| Hardware to self-host | Python 3.10 | Jul 20 | [source](https://github.com/FunAudioLLM/CosyVoice) |
| Project maintenance status | active | Jul 20 | [source](https://github.com/FunAudioLLM/CosyVoice) |
| GitHub stars | 22,288 stars | Jul 20 | [source](https://github.com/FunAudioLLM/CosyVoice) |

## CosyVoice head-to-head

| Comparison | Record | Per use case |
| --- | --- | --- |
| [CosyVoice vs Fish Speech](https://www.versusref.com/tts/cosyvoice-vs-fish-speech/) | won 1 · lost 3 · tied 0 | Audiobooks: lost; Dubbing: lost; Self-Hosted: won; Voice Agents: lost |

Source: https://www.versusref.com/tts/tools/cosyvoice/
