vsref
MiniMax Speech logo

MiniMax Speech Review

Multilingual cloning-first TTS with aggressive pricing

Among the 46 text-to-speech tools we track, MiniMax Speech has the 10th-widest language coverage - a fit for multilingual and localization projects.

From $5/mo

Facts verified Jul 20, 2026Try MiniMax Speech →

If you sign up through links on this page, vsref may earn a commission; programs exist on both sides of most comparisons, and commissions never change verdicts. How we make money

What we know about MiniMax Speech

MiniMax Speech sits in the text-to-speech apis category, where it is multilingual cloning-first TTS with aggressive pricing. We keep this MiniMax Speech profile grounded in primary sources, each fact dated to when we last confirmed it.

On pricing, MiniMax Speech starts at $5 per month for its entry tier. That is the sticker rate: real production cost usually runs higher once you add a language model, a voice provider, and telephony minutes.

On capabilities, MiniMax Speech covers streaming audio output, realtime websocket api, instant voice cloning, emotion / style controls, word-level timestamps, and pronunciation dictionaries, and does not offer professional voice cloning, ssml support, and self-host / on-prem option. Each of those is verified against MiniMax Speech's own docs or dashboard, not marketing copy.

Among the 46 text-to-speech tools in our matrix, MiniMax Speech leads with the 10th-widest language coverage; the fact sheet below has the raw numbers behind that placement.

In total we track 21 verified facts for MiniMax Speech today, each linking the primary source it came from so you can check our work - and vendor claims we have not measured ourselves are labeled as such on the MiniMax Speech fact sheet below.

Reviewed by vsref Editorialfacts verified Jul 20, 2026Methodology →

MiniMax Speech pricing

Usage-priced at $100 per 1M characters (≈ $0.095 per audio-minute), verified Jul 20, 2026 (source).

MiniMax Speech cost at monthly volume tiers
Monthly volume$ / 1M chars$ / audio-minMonthly bill
200K chars/moHobby project$100$0.095$20
2M chars/moProduct feature$100$0.095$200
20M chars/moAt scale$100$0.095$2,000

Effective rates from published pricing; subscription plans resolve to plan fee plus overage. ~950 characters ≈ 1 audio minute. How we compute costs →

Fact sheet

Pricing
Pricing facts
Price per 1M characters (flagship model)100 $/1M charsJul 20
Price per 1M characters (fast model)60 $/1M charsJul 20
Pricing modelhybridJul 20
Cheapest paid plan$5Jul 20
Enterprise / contact-sales thresholdCustom pricing tier above Business ($999/mo)Jul 20
Capabilities
Capabilities facts
TTFB latency (vendor-claimed)~250 msJul 20
Streaming audio output✓ YesJul 20
Realtime websocket API✓ YesJul 20
Instant voice cloning✓ YesJul 20
Professional voice cloning✗ NoJul 20
Minimum audio for voice cloning10 seconds (max 5 minutes; files up to 20 MB, mp3/m4a/wav)Jul 20
Languages supported40 languagesJul 20
Emotion / style controls✓ YesJul 20
SSML support✗ NoJul 20
Word-level timestamps✓ YesJul 20
Pronunciation dictionaries✓ YesJul 20
Compliance & trust
Compliance & trust facts
Self-host / on-prem option✗ NoJul 20
Model weights licenseclosedJul 20
Build experience
Build experience facts
Concurrency on base plan60 RPM (API default for T2A)Jul 20
Output formatsmp3, pcm, flac, wav, pcmu_raw, pcmu_wav, opusJul 20
Max input per requestunder 10,000 characters per request (HTTP T2A)Jul 20

Considering a switch? Best MiniMax Speech alternatives →

STT in this stack

Text-to-speech is half of a voice pipeline: the other half is the speech-to-text that listens. Compare transcription engines on accuracy, streaming latency, and per-minute price: Deepgram, AssemblyAI, GPT-4o Transcribe, or the full speech-to-text comparison →

Voice agents in this stack

The engine is one layer: a voice agent speaks through its text-to-speech engine, but orchestration, telephony, and turn-taking come from the agent platform. Compare the platforms builders pair MiniMax Speech with: Pipecat, OpenAI Realtime API, Twilio ConversationRelay, or the full voice-agent comparison →

Distribute it

Most MiniMax Speech voiceover ends up in short-form video, and publishing that video across TikTok, YouTube, and Instagram is a scheduling problem with real per-channel pricing. Compare the schedulers creators actually run: Buffer, Postiz, Mixpost, or the social media scheduling platforms compared →

Avatar video in this stack

A cloned or bring-your-own MiniMax Speech voice does not have to stay audio-only: AI avatar video platforms lip-sync it onto a talking avatar for finished video. Compare the platforms: HeyGen, Synthesia, Hedra, or the full avatar-video comparison →

MiniMax Speech head-to-head

MiniMax Speech vs Fish Audio →won 2 · lost 4 · tied 0
AudiobookslostContent CreatorswonDeveloperswonDubbinglostSelf-HostedlostVoice Agentslost