vsref
NVIDIA Parakeet / Riva logo

NVIDIA Parakeet / Riva Review

Open-weights, GPU-accelerated self-hosted STT stack

Among the 43 speech-to-text tools we track, NVIDIA Parakeet / Riva has the 31st-widest language coverage.

See pricing

Facts verified Jul 20, 2026Try NVIDIA Parakeet / Riva →

If you sign up through links on this page, vsref may earn a commission; programs exist on both sides of most comparisons, and commissions never change verdicts. How we make money

What we know about NVIDIA Parakeet / Riva

This is our verified profile of NVIDIA Parakeet / Riva, a speech-to-text apis platform - open-weights, GPU-accelerated self-hosted STT stack. Every fact about NVIDIA Parakeet / Riva below carries the source it came from and the day we checked it.

NVIDIA Parakeet / Riva does not publish a public per-minute rate we have been able to verify, so the price figure here stays blank until we can confirm one. We would rather show nothing than a guessed number; if you have a current NVIDIA Parakeet / Riva quote, it is the fastest way for us to close that gap.

On capabilities, NVIDIA Parakeet / Riva covers language auto-detection, word-level timestamps, custom vocabulary / keyterm boosting, speech translation, and self-host / on-prem option, and does not offer websocket streaming api. Each of those is verified against NVIDIA Parakeet / Riva's own docs or dashboard, not marketing copy.

Among the 43 speech-to-text tools in our matrix, NVIDIA Parakeet / Riva leads with the 31st-widest language coverage; the fact sheet below has the raw numbers behind that placement.

In total we track 21 verified facts for NVIDIA Parakeet / Riva today, each linking the primary source it came from so you can check our work - and vendor claims we have not measured ourselves are labeled as such on the NVIDIA Parakeet / Riva fact sheet below.

Reviewed by vsref Editorialfacts verified Jul 20, 2026Methodology →

Fact sheet

Pricing
Pricing facts
Pricing modelhybridJul 20
Free tier quotaFree hosted trial APIs on build.nvidia.comJul 20
Capabilities
Capabilities facts
WER (third-party benchmark)6.43 % WERJul 20
WER (vendor-claimed)~6.34 % WERJul 20
Languages supported25Jul 20
Language auto-detection✓ YesJul 20
Speaker diarization✓ IncludedJul 20
Word-level timestamps✓ YesJul 20
Custom vocabulary / keyterm boosting✓ YesJul 20
Speech translation✓ YesJul 20
Compliance & trust
Compliance & trust facts
Self-host / on-prem option✓ YesJul 20
Model weights licenseCC-BY-4.0Jul 20
Build experience
Build experience facts
Official SDKsPython, Go clients; gRPC protos; CLI clientsJul 20
Max file size / duration24 min full attention; up to 3 hr with local attentionJul 20
Supported audio formatsWAV, FLAC (16 kHz mono); Opus streams in RivaJul 20
Websocket streaming API✗ NoJul 20
Commercial
Commercial facts
Model size (parameters)Parakeet TDT 0.6B (600M); Canary 1B v2 (978M)Jul 20
Hardware to self-hostNVIDIA GPU (T4 to H100); min 2GB RAMJul 20
Hosted API availableFree trial NIM endpoints on build.nvidia.comJul 20
Project maintenance statusActive - NeMo v2.7.3 released 2026-04-23Jul 20
GitHub stars17,800Jul 20

Considering a switch? Best NVIDIA Parakeet / Riva alternatives →

Voice agents in this stack

The engine is one layer: a voice agent hears through its transcription engine, but orchestration, telephony, and turn-taking come from the agent platform. Compare the platforms builders pair NVIDIA Parakeet / Riva with: Pipecat, OpenAI Realtime API, Twilio ConversationRelay, or the full voice-agent comparison →

Want it done for you?

NVIDIA Parakeet / Riva gets you a transcript; the bot that joins the call, labels speakers, and writes the summary is still your build. If that is more pipeline than you want to own, AI meeting notetakers do the whole job end to end: Otter.ai, Fireflies.ai, Fathom, or the full notetaker comparison.

NVIDIA Parakeet / Riva head-to-head

Call CenterswonDeveloperslostDictationwonMedicallostMeetingswonSelf-HostedwonVoice Agentstie
DeveloperslostDictationtieMedicallostSelf-HostedwonVoice Agentslost