# Qwen3-ASR Review

> Qwen3-ASR review: Open-weights multilingual ASR models. Verified pricing, features, and the strongest alternatives in Speech-to-text APIs.

![Qwen3-ASR logo](https://www.versusref.com/logos/qwen3-asr.png)

Open-weights multilingual ASR models

Among the 43 speech-to-text tools we track, Qwen3-ASR has the 25th-widest language coverage.

See pricing · Open source · facts verified Jul 20, 2026

## What we know about Qwen3-ASR

Qwen3-ASR sits in the speech-to-text apis category, where it is open-weights multilingual ASR models. We keep this Qwen3-ASR profile grounded in primary sources, each fact dated to when we last confirmed it.

Qwen3-ASR does not publish a public per-minute rate we have been able to verify, so the price figure here stays blank until we can confirm one. We would rather show nothing than a guessed number; if you have a current Qwen3-ASR quote, it is the fastest way for us to close that gap.

On capabilities, Qwen3-ASR covers language auto-detection, word-level timestamps, self-host / on-prem option, and websocket streaming api, and does not offer speech translation. Each of those is verified against Qwen3-ASR's own docs or dashboard, not marketing copy.

Qwen3-ASR ranks the 25th-widest language coverage of the 43 speech-to-text tools we track, so where it lands for you depends on which of those matters more.

In total we track 22 verified facts for Qwen3-ASR today, each linking the primary source it came from so you can check our work - and vendor claims we have not measured ourselves are labeled as such on the Qwen3-ASR fact sheet below.

## Qwen3-ASR pricing

Published rates: batch $0.0021/min · streaming $0.0054/min, verified Jul 20, 2026 ([source](https://www.alibabacloud.com/help/en/model-studio/model-pricing)).

| Monthly volume | Batch bill | Streaming bill |
| --- | --- | --- |
| 1K min/mo | $2.10 | $5.40 |
| 10K min/mo | $21 | $54 |
| 100K min/mo | $210 | $540 |

Sticker rates only; diarization and PII redaction add-ons price in the stack builder.

## Fact sheet

### Pricing

| Fact | Value | Verified | Source |
| --- | --- | --- | --- |
| Batch price per audio minute | 0.002 $/audio-min | Jul 20 | [source](https://www.alibabacloud.com/help/en/model-studio/model-pricing) |
| Streaming price per audio minute | 0.005 $/audio-min | Jul 20 | [source](https://www.alibabacloud.com/help/en/model-studio/model-pricing) |
| Pricing model | usage | Jul 20 | [source](https://www.alibabacloud.com/help/en/model-studio/model-pricing) |
| Free tier quota | 600 min (10 hrs), intl only | Jul 20 | [source](https://www.alibabacloud.com/help/en/model-studio/model-pricing) |

### Capabilities

| Fact | Value | Verified | Source |
| --- | --- | --- | --- |
| WER (third-party benchmark) | 5.8 % WER | Jul 20 | [source](https://artificialanalysis.ai/speech-to-text) |
| WER (vendor-claimed) | ~1.63 % WER | Jul 20 | [source](https://huggingface.co/Qwen/Qwen3-ASR-1.7B) |
| Languages supported | 52 | Jul 20 | [source](https://github.com/QwenLM/Qwen3-ASR) |
| Language auto-detection | ✓  Yes | Jul 20 | [source](https://huggingface.co/Qwen/Qwen3-ASR-1.7B) |
| Speaker diarization | ✗  Not available | Jul 20 | [source](https://github.com/QwenLM/Qwen3-ASR) |
| PII redaction | ✗  Not available | Jul 20 | [source](https://github.com/QwenLM/Qwen3-ASR) |
| Word-level timestamps | ✓  Yes | Jul 20 | [source](https://github.com/QwenLM/Qwen3-ASR) |
| Speech translation | ✗  No | Jul 20 | [source](https://github.com/QwenLM/Qwen3-ASR) |

### Compliance & trust

| Fact | Value | Verified | Source |
| --- | --- | --- | --- |
| Self-host / on-prem option | ✓  Yes | Jul 20 | [source](https://github.com/QwenLM/Qwen3-ASR) |
| Model weights license | Apache-2.0 | Jul 20 | [source](https://huggingface.co/Qwen/Qwen3-ASR-1.7B) |

### Build experience

| Fact | Value | Verified | Source |
| --- | --- | --- | --- |
| Official SDKs | Python (pip, vLLM, Transformers); DashScope API | Jul 20 | [source](https://github.com/QwenLM/Qwen3-ASR) |
| Max file size / duration | Long audio (toolkit chunking) | Jul 20 | [source](https://github.com/QwenLM/Qwen3-ASR) |
| Websocket streaming API | ✓  Yes | Jul 20 | [source](https://github.com/QwenLM/Qwen3-ASR) |

### Commercial

| Fact | Value | Verified | Source |
| --- | --- | --- | --- |
| Model size (parameters) | 0.6B / 1.7B (+0.6B ForcedAligner) | Jul 20 | [source](https://github.com/QwenLM/Qwen3-ASR) |
| Hardware to self-host | NVIDIA GPU (vLLM / Transformers) | Jul 20 | [source](https://huggingface.co/Qwen/Qwen3-ASR-1.7B) |
| Hosted API available | Yes - Alibaba Model Studio (DashScope) | Jul 20 | [source](https://www.alibabacloud.com/help/en/model-studio/model-pricing) |
| Project maintenance status | Active - released 2026-01 | Jul 20 | [source](https://github.com/QwenLM/Qwen3-ASR) |
| GitHub stars | 3,191 | Jul 20 | [source](https://github.com/QwenLM/Qwen3-ASR) |

## Qwen3-ASR head-to-head

| Comparison | Record | Per use case |
| --- | --- | --- |
| [Qwen3-ASR vs OpenAI Whisper (API)](https://www.versusref.com/stt/qwen3-asr-vs-whisper/) | won 4 · lost 1 · tied 2 | Call Centers: won; Developers: won; Dictation: tie; Medical: lost; Meetings: tie; Self-Hosted: won; Voice Agents: won |

Source: https://www.versusref.com/stt/tools/qwen3-asr/
