vsref
Voxtral TTS logo

Voxtral TTS Review

Open-weights-friendly voice cloning TTS from a frontier AI lab

Among the 15 text-to-speech tools we track, Voxtral TTS has the 1st-fastest time-to-first-byte and the 5th-cheapest flagship rate - a fit for real-time, conversational apps and cost-sensitive, high-volume work.

See pricing

Facts verified Jul 20, 2026Try Voxtral TTS →

If you sign up through links on this page, vsref may earn a commission; programs exist on both sides of most comparisons, and commissions never change verdicts. How we make money

What we know about Voxtral TTS

This is our verified profile of Voxtral TTS, a text-to-speech apis platform - open-weights-friendly voice cloning TTS from a frontier AI lab. Every fact about Voxtral TTS below carries the source it came from and the day we checked it.

Voxtral TTS does not publish a public per-minute rate we have been able to verify, so the price figure here stays blank until we can confirm one. We would rather show nothing than a guessed number; if you have a current Voxtral TTS quote, it is the fastest way for us to close that gap.

On capabilities, Voxtral TTS covers streaming audio output, instant voice cloning, and self-host / on-prem option, and does not offer realtime websocket api, professional voice cloning, and ssml support. Each of those is verified against Voxtral TTS's own docs or dashboard, not marketing copy.

Voxtral TTS ranks the 1st-fastest time-to-first-byte of the 15 text-to-speech tools we track, but only 26th of 15 for languages, so where it lands for you depends on which of those matters more.

In total we track 16 verified facts for Voxtral TTS today, each linking the primary source it came from so you can check our work - and vendor claims we have not measured ourselves are labeled as such on the Voxtral TTS fact sheet below.

Reviewed by vsref Editorialfacts verified Jul 20, 2026Methodology →

Voxtral TTS pricing

Usage-priced at $16 per 1M characters (≈ $0.0152 per audio-minute), verified Jul 20, 2026 (source).

Voxtral TTS cost at monthly volume tiers
Monthly volume$ / 1M chars$ / audio-minMonthly bill
200K chars/moHobby project$16$0.0152$3.20
2M chars/moProduct feature$16$0.0152$32
20M chars/moAt scale$16$0.0152$320

Effective rates from published pricing; subscription plans resolve to plan fee plus overage. ~950 characters ≈ 1 audio minute. How we compute costs →

Fact sheet

Pricing
Pricing facts
Price per 1M characters (flagship model)16 $/1M charsJul 20
Pricing modelusageJul 20
Capabilities
Capabilities facts
TTFB latency (vendor-claimed)~70 msJul 20
Streaming audio output✓ YesJul 20
Realtime websocket API✗ NoJul 20
Instant voice cloning✓ YesJul 20
Professional voice cloning✗ NoJul 20
Minimum audio for voice cloning~3 seconds (docs say 2-3 seconds)Jul 20
Languages supported9 languagesJul 20
SSML support✗ NoJul 20
Word-level timestamps✗ NoJul 20
Pronunciation dictionaries✗ NoJul 20
Compliance & trust
Compliance & trust facts
Self-host / on-prem option✓ YesJul 20
Model weights licenseCC BY-NC 4.0 (open weights, non-commercial)Jul 20
Build experience
Build experience facts
Official SDKsOfficial Python (mistralai) and TypeScript SDKsJul 20
Output formatspcm, mp3Jul 20

Considering a switch? Best Voxtral TTS alternatives →

STT in this stack

Text-to-speech is half of a voice pipeline: the other half is the speech-to-text that listens. Compare transcription engines on accuracy, streaming latency, and per-minute price: Deepgram, AssemblyAI, GPT-4o Transcribe, or the full speech-to-text comparison →

Voice agents in this stack

The engine is one layer: a voice agent speaks through its text-to-speech engine, but orchestration, telephony, and turn-taking come from the agent platform. Compare the platforms builders pair Voxtral TTS with: Pipecat, OpenAI Realtime API, Twilio ConversationRelay, or the full voice-agent comparison →

Distribute it

Most Voxtral TTS voiceover ends up in short-form video, and publishing that video across TikTok, YouTube, and Instagram is a scheduling problem with real per-channel pricing. Compare the schedulers creators actually run: Buffer, Postiz, Mixpost, or the social media scheduling platforms compared →

Avatar video in this stack

A cloned or bring-your-own Voxtral TTS voice does not have to stay audio-only: AI avatar video platforms lip-sync it onto a talking avatar for finished video. Compare the platforms: HeyGen, Synthesia, Hedra, or the full avatar-video comparison →

Voxtral TTS head-to-head

Voxtral TTS vs ElevenLabs →won 1 · lost 3 · tied 0
DeveloperslostDubbinglostSelf-HostedwonVoice Agentslost
Voxtral TTS vs Cartesia →won 1 · lost 5 · tied 0
AudiobookslostContent CreatorslostDeveloperslostDubbinglostSelf-HostedwonVoice Agentslost
Voxtral TTS vs OpenAI TTS →won 3 · lost 3 · tied 0
AudiobookswonContent CreatorslostDeveloperslostDubbingwonSelf-HostedwonVoice Agentslost