# Voxtral TTS Review

> Voxtral TTS review: Open-weights-friendly voice cloning TTS from a frontier AI lab. Verified pricing, features, and the strongest alternatives in.

![Voxtral TTS logo](https://www.versusref.com/logos/voxtral-tts.png)

Open-weights-friendly voice cloning TTS from a frontier AI lab

Among the 15 text-to-speech tools we track, Voxtral TTS has the 1st-fastest time-to-first-byte and the 5th-cheapest flagship rate - a fit for real-time, conversational apps and cost-sensitive, high-volume work.

See pricing · facts verified Jul 20, 2026

## What we know about Voxtral TTS

This is our verified profile of Voxtral TTS, a text-to-speech apis platform - open-weights-friendly voice cloning TTS from a frontier AI lab. Every fact about Voxtral TTS below carries the source it came from and the day we checked it.

Voxtral TTS does not publish a public per-minute rate we have been able to verify, so the price figure here stays blank until we can confirm one. We would rather show nothing than a guessed number; if you have a current Voxtral TTS quote, it is the fastest way for us to close that gap.

On capabilities, Voxtral TTS covers streaming audio output, instant voice cloning, and self-host / on-prem option, and does not offer realtime websocket api, professional voice cloning, and ssml support. Each of those is verified against Voxtral TTS's own docs or dashboard, not marketing copy.

Voxtral TTS ranks the 1st-fastest time-to-first-byte of the 15 text-to-speech tools we track, but only 26th of 15 for languages, so where it lands for you depends on which of those matters more.

In total we track 16 verified facts for Voxtral TTS today, each linking the primary source it came from so you can check our work - and vendor claims we have not measured ourselves are labeled as such on the Voxtral TTS fact sheet below.

## Voxtral TTS pricing

Usage-priced at $16 per 1M characters (≈ $0.0152 per audio-minute), verified Jul 20, 2026 ([source](https://mistral.ai/news/voxtral-tts)).

| Monthly volume | $ / 1M chars | $ / audio-min | Monthly bill |
| --- | --- | --- | --- |
| 200K chars/mo | $16 | $0.0152 | $3.20 |
| 2M chars/mo | $16 | $0.0152 | $32 |
| 20M chars/mo | $16 | $0.0152 | $320 |

Effective rates from published pricing; subscription plans resolve to plan fee plus overage. ~950 characters ≈ 1 audio minute.

## Fact sheet

### Pricing

| Fact | Value | Verified | Source |
| --- | --- | --- | --- |
| Price per 1M characters (flagship model) | 16 $/1M chars | Jul 20 | [source](https://mistral.ai/news/voxtral-tts) |
| Pricing model | usage | Jul 20 | [source](https://mistral.ai/news/voxtral-tts) |

### Capabilities

| Fact | Value | Verified | Source |
| --- | --- | --- | --- |
| TTFB latency (vendor-claimed) | ~70 ms | Jul 20 | [source](https://mistral.ai/news/voxtral-tts) |
| Streaming audio output | ✓  Yes | Jul 20 | [source](https://docs.mistral.ai/capabilities/audio/) |
| Realtime websocket API | ✗  No | Jul 20 | [source](https://docs.mistral.ai/studio-api/audio/text_to_speech) |
| Instant voice cloning | ✓  Yes | Jul 20 | [source](https://mistral.ai/news/voxtral-tts) |
| Professional voice cloning | ✗  No | Jul 20 | [source](https://mistral.ai/news/voxtral-tts) |
| Minimum audio for voice cloning | ~3 seconds (docs say 2-3 seconds) | Jul 20 | [source](https://mistral.ai/news/voxtral-tts) |
| Languages supported | 9 languages | Jul 20 | [source](https://mistral.ai/news/voxtral-tts) |
| SSML support | ✗  No | Jul 20 | [source](https://docs.mistral.ai/studio-api/audio/text_to_speech) |
| Word-level timestamps | ✗  No | Jul 20 | [source](https://docs.mistral.ai/studio-api/audio/text_to_speech) |
| Pronunciation dictionaries | ✗  No | Jul 20 | [source](https://docs.mistral.ai/studio-api/audio/text_to_speech) |

### Compliance & trust

| Fact | Value | Verified | Source |
| --- | --- | --- | --- |
| Self-host / on-prem option | ✓  Yes | Jul 20 | [source](https://mistral.ai/news/voxtral-tts) |
| Model weights license | CC BY-NC 4.0 (open weights, non-commercial) | Jul 20 | [source](https://mistral.ai/news/voxtral-tts) |

### Build experience

| Fact | Value | Verified | Source |
| --- | --- | --- | --- |
| Official SDKs | Official Python (mistralai) and TypeScript SDKs | Jul 20 | [source](https://docs.mistral.ai/getting-started/clients/) |
| Output formats | pcm, mp3 | Jul 20 | [source](https://docs.mistral.ai/studio-api/audio/text_to_speech) |

## Voxtral TTS head-to-head

| Comparison | Record | Per use case |
| --- | --- | --- |
| [Voxtral TTS vs ElevenLabs](https://www.versusref.com/tts/elevenlabs-vs-voxtral-tts/) | won 1 · lost 3 · tied 0 | Developers: lost; Dubbing: lost; Self-Hosted: won; Voice Agents: lost |
| [Voxtral TTS vs Cartesia](https://www.versusref.com/tts/cartesia-vs-voxtral-tts/) | won 1 · lost 5 · tied 0 | Audiobooks: lost; Content Creators: lost; Developers: lost; Dubbing: lost; Self-Hosted: won; Voice Agents: lost |
| [Voxtral TTS vs OpenAI TTS](https://www.versusref.com/tts/openai-tts-vs-voxtral-tts/) | won 3 · lost 3 · tied 0 | Audiobooks: won; Content Creators: lost; Developers: lost; Dubbing: won; Self-Hosted: won; Voice Agents: lost |

Source: https://www.versusref.com/tts/tools/voxtral-tts/
