# Fish Audio vs MiniMax Speech

> Fish Audio vs MiniMax Speech: Fish Audio is our pick, leading for audiobooks and dubbing. They split on concurrency on base plan and output formats. Verified.

![Fish Audio logo](https://www.versusref.com/logos/fish-audio.png) ![MiniMax Speech logo](https://www.versusref.com/logos/minimax-speech.png)

This is the verified Fish Audio vs MiniMax Speech breakdown: Fish Audio (Developer-first hosted API from the team behind the open-source fish-speech models; simple prepaid pay-as-you-go billing, 80+ language coverage, and a free fair-use model tier (s2.1-pro-free). Facts here cover the hosted API, not the OSS model.) and MiniMax Speech (Multilingual cloning-first TTS with aggressive pricing) are Text-to-speech APIs platforms, priced from $15/1M chars and $100/1M chars respectively. The bottom line, head-to-head facts, and true cost tiers follow.

Text-to-speech APIs platforms · 23 facts compared · all sourced · pricing verified Jul 20, 2026

## Bottom line

Fish Audio wins 4 of 6 use case verdicts versus MiniMax Speech's 2. It also leads on core metrics: flagship pricing at 15 $/1M UTF-8 bytes versus 100 $/1M chars, TTFB of 100 ms versus 250 ms, 83 languages versus 40, and a self-host option MiniMax lacks entirely. MiniMax retains advantages in concurrency (60 RPM base vs 5) and developer tooling extras like word timestamps. Even so, Fish Audio is the safer default for most buyers across price, latency, and deployment flexibility.

## Verdict summary

| Use case | Winner | Why |
| --- | --- | --- |
| Audiobooks | Fish Audio | For audiobook production, per-character cost is decisive: Fish Audio charges $15 per 1M UTF-8 bytes versus MiniMax Speech at $100 per 1M characters, making Fish Audio roughly 6x cheaper at scale. |
| Content Creators | MiniMax Speech | For content creators, MiniMax Speech offers a predictable hybrid pricing model with a $5/mo entry plan and tiered subscriptions, matching the need for budget predictability. |
| Developers | MiniMax Speech | For developers metering usage, MiniMax Speech offers word-level timestamps, pronunciation dictionaries, and 7 output formats including flac, opus, and pcmu variants. |
| Dubbing | Fish Audio | For dubbing and localization, language coverage is the primary deciding factor. |
| Self-Hosted | Fish Audio | Fish Audio explicitly supports a self-host/on-prem option, while MiniMax Speech offers no self-host option at all. |
| Voice Agents | Fish Audio | For voice agents, latency is decisive. |

## Fish Audio vs MiniMax Speech: head-to-head facts

| Fact | Fish Audio | MiniMax Speech | Verified |
| --- | --- | --- | --- |
| Price per 1M characters (flagship model) | 15 $/1M UTF-8 bytes ([source](https://docs.fish.audio/developer-guide/models-pricing/pricing-and-rate-limits)) | 100 $/1M chars ([source](https://platform.minimax.io/docs/guides/pricing-paygo.md)) | Jul 20 |
| Cheapest paid plan | n/a | $5 ([source](https://platform.minimax.io/docs/guides/pricing-speech.md)) | Jul 20 |
| Free tier quota | s2.1-pro-free model ([source](https://docs.fish.audio/developer-guide/models-pricing/pricing-and-rate-limits)) | n/a | Jul 20 |
| TTFB latency (vendor-claimed) | ~100 ms ([source](https://docs.fish.audio/developer-guide/models-pricing/models-overview)) | ~250 ms ([source](https://www.minimax.io/news/minimax-speech-26)) | Jul 20 |
| Streaming audio output | ✓  Yes ([source](https://docs.fish.audio/features/realtime-streaming)) | ✓  Yes ([source](https://platform.minimax.io/docs/api-reference/speech-t2a-http.md)) | Jul 20 |
| Instant voice cloning | ✓  Yes ([source](https://docs.fish.audio/features/voice-cloning)) | ✓  Yes ([source](https://platform.minimax.io/docs/guides/pricing-paygo.md)) | Jul 20 |
| Languages supported | ~83 languages (s2.1-pro, auto-detected) ([source](https://docs.fish.audio/developer-guide/models-pricing/models-overview)) | 40 languages ([source](https://platform.minimax.io/docs/api-reference/speech-t2a-http.md)) | Jul 20 |
| Model weights license | n/a | closed ([source](https://platform.minimax.io/docs/api-reference/speech-t2a-http.md)) | Jul 20 |

## Fish Audio vs MiniMax Speech pricing: true cost at 3 usage tiers

Effective monthly bill from published rates. Subscription plans resolve to plan fee plus overage; ~950 characters ≈ 1 audio minute.

| Volume | Buyer | Fish Audio | MiniMax Speech | Note |
| --- | --- | --- | --- | --- |
| 200K chars/mo | Hobby project | $3 | $20 | About 3.5 hours of audio. Free tiers may cover part of this. |
| 2M chars/mo | Product feature | $30 | $200 | About 35 hours of audio a month. |
| 20M chars/mo | At scale | $300 | $2,000 | About 350 hours of audio. Most vendors negotiate at this tier. |

## Fish Audio vs MiniMax Speech: verdicts by use case

### Audiobooks → Fish Audio

For audiobook production, per-character cost is decisive: Fish Audio charges $15 per 1M UTF-8 bytes versus MiniMax Speech at $100 per 1M characters, making Fish Audio roughly 6x cheaper at scale. MiniMax also caps input at under 10,000 characters per request, which is a significant friction point for book-length content. Fish Audio additionally supports 83 languages versus 40 for MiniMax, and offers a self-host option for studios wanting on-premises control. MiniMax has pronunciation dictionaries, but cost and input-length constraints outweigh that advantage for long-form narration.

### Content Creators → MiniMax Speech

For content creators, MiniMax Speech offers a predictable hybrid pricing model with a $5/mo entry plan and tiered subscriptions, matching the need for budget predictability. It supports 40 languages with word-level timestamps, pronunciation dictionaries, and emotion and style controls, all useful for polished voiceover work. Fish Audio is cheaper per character at $15/1M bytes versus $100/1M chars for MiniMax flagship, but MiniMax's subscription structure and richer production features like timestamps and pronunciation dictionaries tip it narrowly for content creator workflows where consistency and tooling matter more than raw API cost.

### Developers → MiniMax Speech

For developers metering usage, MiniMax Speech offers word-level timestamps, pronunciation dictionaries, and 7 output formats including flac, opus, and pcmu variants. Fish Audio lacks verified timestamp support and charges $15 per 1M UTF-8 bytes, compared to MiniMax at $60-100 per 1M characters. MiniMax hybrid pricing also adds a $5 per month platform fee. Fish Audio wins on TTFB (100ms vs 250ms) and self-hosting. Overall, MiniMax's timestamps and richer output formats tip the developer tooling comparison narrowly in its favor.

### Dubbing → Fish Audio

For dubbing and localization, language coverage is the primary deciding factor. Fish Audio supports 83 languages versus MiniMax Speech's 40 languages, giving Fish Audio more than double the reach. Both tools offer instant voice cloning with a 10-second minimum, so cross-language voice cloning capability is roughly equivalent. Fish Audio also costs significantly less at 15 dollars per 1M UTF-8 bytes versus 100 dollars per 1M chars for MiniMax, which matters when scaling across many language versions of the same content.

### Self-Hosted → Fish Audio

Fish Audio explicitly supports a self-host/on-prem option, while MiniMax Speech offers no self-host option at all. MiniMax also uses closed model weights, making any on-prem deployment impossible regardless. Fish Audio wins decisively on both the deployment flexibility and model openness dimensions that define this use case.

### Voice Agents → Fish Audio

For voice agents, latency is decisive. Fish Audio claims a 100 ms TTFB versus MiniMax's 250 ms, a 2.5x advantage that directly determines conversational naturalness. Both support real-time WebSocket APIs and streaming output. Fish Audio also offers 83 languages versus 40 for MiniMax, which is useful for multilingual agent deployments. Its pure usage pricing without a platform fee is simpler at lower volumes. The margin is narrow because these TTFB figures are vendor-claimed and not independently verified, but the gap is large enough to favor Fish Audio.

## Fish Audio vs MiniMax Speech: common questions

### Which is cheaper for high-volume text to speech, MiniMax Speech or Fish Audio?

Fish Audio charges 15 dollars per 1M UTF-8 bytes for its flagship model. MiniMax Speech charges 100 dollars per 1M characters for its flagship model and 60 dollars per 1M characters for its fast model. Fish Audio is significantly cheaper for large-scale usage at those rates, though the two services measure units differently (bytes vs. characters), so direct comparisons require that caveat.

### Does Fish Audio have a free tier?

Yes. Fish Audio offers a free tier that includes access to the s2.1-pro-free model, so you can test the service without paying. MiniMax Speech uses a hybrid pricing model with a cheapest paid plan starting at 5 dollars per month.

### Which service has lower latency for real-time voice applications?

According to vendor-claimed figures, Fish Audio reports a TTFB of 100 ms while MiniMax Speech reports 250 ms. Both services support a real-time WebSocket API and streaming audio output. Note that these latency figures are vendor-claimed and have not been independently verified.

### Can I clone a voice with either MiniMax Speech or Fish Audio?

Both tools support instant voice cloning with a minimum of 10 seconds of audio. MiniMax Speech accepts files up to 20 MB in mp3, m4a, or wav format, with a maximum duration of 5 minutes. However, MiniMax Speech does not offer professional voice cloning, while Fish Audio lists its cloning minimum simply as 10 seconds.

### How many languages does each service support?

Fish Audio supports 83 languages with auto-detection on its s2.1-pro model (vendor-claimed), while MiniMax Speech supports 40 languages (verified). Fish Audio therefore covers a broader language range, though its count comes from the vendor rather than independent verification.

### Can I self-host or deploy Fish Audio or MiniMax Speech on my own infrastructure?

Fish Audio offers a self-host and on-premises option (vendor-claimed), while MiniMax Speech does not offer self-host or on-premises deployment, and its model weights are closed. If data residency or private deployment is a requirement, Fish Audio is the better choice.

### What audio output formats do MiniMax Speech and Fish Audio support?

Fish Audio outputs mp3 (default), wav, pcm, and opus. MiniMax Speech outputs mp3, pcm, flac, wav, pcmu_raw, pcmu_wav, and opus. MiniMax Speech supports a wider format selection, including flac and pcmu variants, which may be important for telephony integrations.

### Does MiniMax Speech support SSML for fine-grained speech control?

MiniMax Speech does not support SSML. It instead offers emotion and style controls, word-level timestamps, and pronunciation dictionaries as alternative ways to control output. Fish Audio also offers emotion and style controls, but whether it supports SSML is not documented in the available facts.

Source: https://www.versusref.com/tts/fish-audio-vs-minimax-speech/
