vsref
NVIDIA Magpie TTS logo

NVIDIA Magpie TTS Review

OSS

A small, GPU-efficient 9-language TTS checkpoint for teams already in the NVIDIA NeMo/Riva ecosystem; commercially usable open weights, but zero-shot voice cloning was removed from the open release and it caps generations at about 20 seconds.

Among the 46 text-to-speech tools we track, NVIDIA Magpie TTS has the 26th-widest language coverage.

See pricing

Facts verified Jul 20, 2026Website →

What we know about NVIDIA Magpie TTS

This is our verified profile of NVIDIA Magpie TTS, a text-to-speech apis platform - a small, GPU-efficient 9-language TTS checkpoint for teams already in the NVIDIA NeMo/Riva ecosystem; commercially usable open weights, but zero-shot voice cloning was removed from the open release and it caps generations at about 20 seconds. Every fact about NVIDIA Magpie TTS below carries the source it came from and the day we checked it.

NVIDIA Magpie TTS does not publish a public per-minute rate we have been able to verify, so the price figure here stays blank until we can confirm one. We would rather show nothing than a guessed number; if you have a current NVIDIA Magpie TTS quote, it is the fastest way for us to close that gap.

On capabilities, NVIDIA Magpie TTS covers self-host / on-prem option, and does not offer instant voice cloning. Each of those is verified against NVIDIA Magpie TTS's own docs or dashboard, not marketing copy.

Among the 46 text-to-speech tools in our matrix, NVIDIA Magpie TTS leads with the 26th-widest language coverage; the fact sheet below has the raw numbers behind that placement.

In total we track 11 verified facts for NVIDIA Magpie TTS today, each linking the primary source it came from so you can check our work - and vendor claims we have not measured ourselves are labeled as such on the NVIDIA Magpie TTS fact sheet below.

Reviewed by vsref Editorialfacts verified Jul 20, 2026Methodology →

Fact sheet

Capabilities
Capabilities facts
Instant voice cloning✗ NoJul 20
Languages supported9 languagesJul 20
Compliance & trust
Compliance & trust facts
Self-host / on-prem option✓ YesJul 20
Model weights licenseNVIDIA Open Model License AgreementJul 20
Build experience
Build experience facts
Official SDKsNVIDIA NeMo Framework (Python) for local inferenceJul 20
Output formats22.05 kHz audio (neural codec output)Jul 20
Max input per requestUp to 20s speech per requestJul 20
Commercial
Commercial facts
Model size (parameters)357MJul 20
Hardware to self-hostNVIDIA A10, A30, A100, or H100 GPU; Linux preferredJul 20
Hosted API availableYesJul 20
Project maintenance statusActiveJul 20

Considering a switch? Best NVIDIA Magpie TTS alternatives →

STT in this stack

Text-to-speech is half of a voice pipeline: the other half is the speech-to-text that listens. Compare transcription engines on accuracy, streaming latency, and per-minute price: Deepgram, AssemblyAI, GPT-4o Transcribe, or the full speech-to-text comparison →

Voice agents in this stack

The engine is one layer: a voice agent speaks through its text-to-speech engine, but orchestration, telephony, and turn-taking come from the agent platform. Compare the platforms builders pair NVIDIA Magpie TTS with: Pipecat, OpenAI Realtime API, Twilio ConversationRelay, or the full voice-agent comparison →

Distribute it

Most NVIDIA Magpie TTS voiceover ends up in short-form video, and publishing that video across TikTok, YouTube, and Instagram is a scheduling problem with real per-channel pricing. Compare the schedulers creators actually run: Buffer, Postiz, Mixpost, or the social media scheduling platforms compared →

Avatar video in this stack

A cloned or bring-your-own NVIDIA Magpie TTS voice does not have to stay audio-only: AI avatar video platforms lip-sync it onto a talking avatar for finished video. Compare the platforms: HeyGen, Synthesia, Hedra, or the full avatar-video comparison →