# Voxtral (open weights) Review

> Voxtral (open weights) review: A frontier-lab open-weight TTS you can run on a single 16GB GPU; the CC BY-NC license makes it evaluation/research-only, with.

![Voxtral (open weights) logo](https://www.versusref.com/logos/voxtral-open.png)

A frontier-lab open-weight TTS you can run on a single 16GB GPU; the CC BY-NC license makes it evaluation/research-only, with Mistral's paid API as the commercial route.

Among the 46 text-to-speech tools we track, Voxtral (open weights) has the 26th-widest language coverage.

See pricing · Open source · facts verified Jul 20, 2026

## What we know about Voxtral (open weights)

This is our verified profile of Voxtral (open weights), a text-to-speech apis platform - a frontier-lab open-weight TTS you can run on a single 16GB GPU; the CC BY-NC license makes it evaluation/research-only, with Mistral's paid API as the commercial route. Every fact about Voxtral (open weights) below carries the source it came from and the day we checked it.

Voxtral (open weights) does not publish a public per-minute rate we have been able to verify, so the price figure here stays blank until we can confirm one. We would rather show nothing than a guessed number; if you have a current Voxtral (open weights) quote, it is the fastest way for us to close that gap.

On capabilities, Voxtral (open weights) covers streaming audio output, instant voice cloning, emotion / style controls, and self-host / on-prem option. Each of those is verified against Voxtral (open weights)'s own docs or dashboard, not marketing copy.

Placed against the 46 text-to-speech tools we track, Voxtral (open weights)'s strongest showing is the 26th-widest language coverage - a spread worth weighing against your own priorities.

In total we track 12 verified facts for Voxtral (open weights) today, each linking the primary source it came from so you can check our work - and vendor claims we have not measured ourselves are labeled as such on the Voxtral (open weights) fact sheet below.

## Fact sheet

### Capabilities

| Fact | Value | Verified | Source |
| --- | --- | --- | --- |
| Streaming audio output | ✓  Yes | Jul 20 | [source](https://huggingface.co/mistralai/Voxtral-4B-TTS-2603) |
| Instant voice cloning | ✓  Yes | Jul 20 | [source](https://huggingface.co/mistralai/Voxtral-4B-TTS-2603) |
| Languages supported | 9 languages | Jul 20 | [source](https://huggingface.co/mistralai/Voxtral-4B-TTS-2603) |
| Emotion / style controls | ✓  Yes | Jul 20 | [source](https://huggingface.co/mistralai/Voxtral-4B-TTS-2603) |

### Compliance & trust

| Fact | Value | Verified | Source |
| --- | --- | --- | --- |
| Self-host / on-prem option | ✓  Yes | Jul 20 | [source](https://huggingface.co/mistralai/Voxtral-4B-TTS-2603) |
| Model weights license | CC-BY-NC-4.0 | Jul 20 | [source](https://huggingface.co/mistralai/Voxtral-4B-TTS-2603) |

### Build experience

| Fact | Value | Verified | Source |
| --- | --- | --- | --- |
| Official SDKs | vLLM-Omni (>=0.18.0 recommended) | Jul 20 | [source](https://huggingface.co/mistralai/Voxtral-4B-TTS-2603) |
| Output formats | WAV, PCM, FLAC, MP3, AAC, Opus (24 kHz) | Jul 20 | [source](https://huggingface.co/mistralai/Voxtral-4B-TTS-2603) |

### Commercial

| Fact | Value | Verified | Source |
| --- | --- | --- | --- |
| Model size (parameters) | 4B | Jul 20 | [source](https://huggingface.co/mistralai/Voxtral-4B-TTS-2603) |
| Hardware to self-host | Minimum 16GB GPU VRAM (single GPU) | Jul 20 | [source](https://huggingface.co/mistralai/Voxtral-4B-TTS-2603) |
| Hosted API available | Yes | Jul 20 | [source](https://mistral.ai/news/voxtral-tts/) |
| Project maintenance status | active | Jul 20 | [source](https://mistral.ai/news/voxtral-tts/) |

Source: https://www.versusref.com/tts/tools/voxtral-open/
