vsref
Fish Speech logo

Fish Speech Review

OSS

Top-tier expressive multilingual open-weights TTS whose license moved from permissive to research/non-commercial; commercial use requires a license from Fish Audio or their hosted API.

Among the 46 text-to-speech tools we track, Fish Speech has the 6th-widest language coverage - a fit for multilingual and localization projects.

See pricing

Facts verified Jul 20, 2026Website →

What we know about Fish Speech

This is our verified profile of Fish Speech, a text-to-speech apis platform - top-tier expressive multilingual open-weights TTS whose license moved from permissive to research/non-commercial; commercial use requires a license from Fish Audio or their hosted API. Every fact about Fish Speech below carries the source it came from and the day we checked it.

Fish Speech does not publish a public per-minute rate we have been able to verify, so the price figure here stays blank until we can confirm one. We would rather show nothing than a guessed number; if you have a current Fish Speech quote, it is the fastest way for us to close that gap.

On capabilities, Fish Speech covers streaming audio output, instant voice cloning, emotion / style controls, and self-host / on-prem option. Each of those is verified against Fish Speech's own docs or dashboard, not marketing copy.

Fish Speech ranks the 6th-widest language coverage of the 46 text-to-speech tools we track, so where it lands for you depends on which of those matters more.

In total we track 12 verified facts for Fish Speech today, each linking the primary source it came from so you can check our work - and vendor claims we have not measured ourselves are labeled as such on the Fish Speech fact sheet below.

Reviewed by vsref Editorialfacts verified Jul 20, 2026Methodology →

Fact sheet

Capabilities
Capabilities facts
Streaming audio output✓ YesJul 20
Instant voice cloning✓ YesJul 20
Languages supported80 languagesJul 20
Emotion / style controls✓ YesJul 20
Compliance & trust
Compliance & trust facts
Self-host / on-prem option✓ YesJul 20
Model weights licenseFish Audio Research License (custom, non-commercial)Jul 20
Build experience
Build experience facts
Official SDKsPyTorch; vLLM (Omni) and SGLang serving; DockerJul 20
Commercial
Commercial facts
Model size (parameters)4B (current flagship, S2 Pro per READMEJul 20
Hardware to self-hostGPU recommended (benchmarks on NVIDIA H200)Jul 20
Hosted API availableYesJul 20
Project maintenance statusactiveJul 20
GitHub stars31,300 starsJul 20

Considering a switch? Best Fish Speech alternatives →

STT in this stack

Text-to-speech is half of a voice pipeline: the other half is the speech-to-text that listens. Compare transcription engines on accuracy, streaming latency, and per-minute price: Deepgram, AssemblyAI, GPT-4o Transcribe, or the full speech-to-text comparison →

Voice agents in this stack

The engine is one layer: a voice agent speaks through its text-to-speech engine, but orchestration, telephony, and turn-taking come from the agent platform. Compare the platforms builders pair Fish Speech with: Pipecat, OpenAI Realtime API, Twilio ConversationRelay, or the full voice-agent comparison →

Distribute it

Most Fish Speech voiceover ends up in short-form video, and publishing that video across TikTok, YouTube, and Instagram is a scheduling problem with real per-channel pricing. Compare the schedulers creators actually run: Buffer, Postiz, Mixpost, or the social media scheduling platforms compared →

Avatar video in this stack

A cloned or bring-your-own Fish Speech voice does not have to stay audio-only: AI avatar video platforms lip-sync it onto a talking avatar for finished video. Compare the platforms: HeyGen, Synthesia, Hedra, or the full avatar-video comparison →

Fish Speech head-to-head

Fish Speech vs CosyVoice →won 3 · lost 1 · tied 0
AudiobookswonDubbingwonSelf-HostedlostVoice Agentswon