vsref
Qwen3-ASR logo

Qwen3-ASR Review

OSS

Open-weights multilingual ASR models

Among the 43 speech-to-text tools we track, Qwen3-ASR has the 25th-widest language coverage.

See pricing

Facts verified Jul 20, 2026Website →

What we know about Qwen3-ASR

Qwen3-ASR sits in the speech-to-text apis category, where it is open-weights multilingual ASR models. We keep this Qwen3-ASR profile grounded in primary sources, each fact dated to when we last confirmed it.

Qwen3-ASR does not publish a public per-minute rate we have been able to verify, so the price figure here stays blank until we can confirm one. We would rather show nothing than a guessed number; if you have a current Qwen3-ASR quote, it is the fastest way for us to close that gap.

On capabilities, Qwen3-ASR covers language auto-detection, word-level timestamps, self-host / on-prem option, and websocket streaming api, and does not offer speech translation. Each of those is verified against Qwen3-ASR's own docs or dashboard, not marketing copy.

Qwen3-ASR ranks the 25th-widest language coverage of the 43 speech-to-text tools we track, so where it lands for you depends on which of those matters more.

In total we track 22 verified facts for Qwen3-ASR today, each linking the primary source it came from so you can check our work - and vendor claims we have not measured ourselves are labeled as such on the Qwen3-ASR fact sheet below.

Reviewed by vsref Editorialfacts verified Jul 20, 2026Methodology →

Qwen3-ASR pricing

Published rates: batch $0.0021/min · streaming $0.0054/min, verified Jul 20, 2026 (source).

Qwen3-ASR cost at monthly volume tiers
Monthly volumeBatch billStreaming bill
1K min/moSide project$2.10$5.40
10K min/moProduction app$21$54
100K min/moCall-center scale$210$540

Sticker rates only; diarization and PII redaction add-ons price in the stack builder. How we compute costs →

Fact sheet

Pricing
Pricing facts
Batch price per audio minute0.002 $/audio-minJul 20
Streaming price per audio minute0.005 $/audio-minJul 20
Pricing modelusageJul 20
Free tier quota600 min (10 hrs), intl onlyJul 20
Capabilities
Capabilities facts
WER (third-party benchmark)5.8 % WERJul 20
WER (vendor-claimed)~1.63 % WERJul 20
Languages supported52Jul 20
Language auto-detection✓ YesJul 20
Speaker diarization✗ Not availableJul 20
PII redaction✗ Not availableJul 20
Word-level timestamps✓ YesJul 20
Speech translation✗ NoJul 20
Compliance & trust
Compliance & trust facts
Self-host / on-prem option✓ YesJul 20
Model weights licenseApache-2.0Jul 20
Build experience
Build experience facts
Official SDKsPython (pip, vLLM, Transformers); DashScope APIJul 20
Max file size / durationLong audio (toolkit chunking)Jul 20
Websocket streaming API✓ YesJul 20
Commercial
Commercial facts
Model size (parameters)0.6B / 1.7B (+0.6B ForcedAligner)Jul 20
Hardware to self-hostNVIDIA GPU (vLLM / Transformers)Jul 20
Hosted API availableYes - Alibaba Model Studio (DashScope)Jul 20
Project maintenance statusActive - released 2026-01Jul 20
GitHub stars3,191Jul 20

Considering a switch? Best Qwen3-ASR alternatives →

Voice agents in this stack

The engine is one layer: a voice agent hears through its transcription engine, but orchestration, telephony, and turn-taking come from the agent platform. Compare the platforms builders pair Qwen3-ASR with: Pipecat, OpenAI Realtime API, Twilio ConversationRelay, or the full voice-agent comparison →

Want it done for you?

Qwen3-ASR gets you a transcript; the bot that joins the call, labels speakers, and writes the summary is still your build. If that is more pipeline than you want to own, AI meeting notetakers do the whole job end to end: Otter.ai, Fireflies.ai, Fathom, or the full notetaker comparison.

Qwen3-ASR head-to-head

Qwen3-ASR vs OpenAI Whisper (API) →won 4 · lost 1 · tied 2
Call CenterswonDeveloperswonDictationtieMedicallostMeetingstieSelf-HostedwonVoice Agentswon