Qwen3-TTS
OSSGenuinely open weights (confirmed - not API-only since the Jan 2026 release): Apache-2.0 checkpoints on Hugging Face with an Alibaba Cloud DashScope API for hosted use.
Among the 46 text-to-speech tools we track, Qwen3-TTS has the 25th-widest language coverage.
See pricing
What we know about Qwen3-TTS
Qwen3-TTS is a text-to-speech apis platform: genuinely open weights (confirmed - not API-only since the Jan 2026 release): Apache-2.0 checkpoints on Hugging Face with an Alibaba Cloud DashScope API for hosted use. This profile tracks what we have verified about Qwen3-TTS, each fact linked to a primary source and dated.
Qwen3-TTS does not publish a public per-minute rate we have been able to verify, so the price figure here stays blank until we can confirm one. We would rather show nothing than a guessed number; if you have a current Qwen3-TTS quote, it is the fastest way for us to close that gap.
On capabilities, Qwen3-TTS covers streaming audio output, instant voice cloning, emotion / style controls, and self-host / on-prem option. Each of those is verified against Qwen3-TTS's own docs or dashboard, not marketing copy.
In total we track 11 verified facts for Qwen3-TTS today, and add coverage as the ingestion pass revisits it. Where a number is a vendor claim rather than our own measurement, it is labeled as such on the Qwen3-TTS fact sheet below, and each row links the primary source it came from so you can check our work.
Whether Qwen3-TTS is the right call depends on your use case more than any single spec, which is why the head-to-head verdicts below score it per scenario rather than crowning one overall winner. Use the fact sheet for the raw numbers, and the matchups for how Qwen3-TTS actually fares against the platforms buyers most often weigh it against.
Fact sheet
Every row independently verifiedConsidering a switch? Best Qwen3-TTS alternatives →