Fish Audio Review
Developer-first hosted API from the team behind the open-source fish-speech models; simple prepaid pay-as-you-go billing, 80+ language coverage, and a free fair-use model tier (s2.1-pro-free). Facts here cover the hosted API, not the OSS model.
Among the 46 text-to-speech tools we track, Fish Audio has the 5th-widest language coverage and the 3rd-cheapest flagship rate - a fit for multilingual and localization projects and cost-sensitive, high-volume work.
See pricing
If you sign up through links on this page, vsref may earn a commission; programs exist on both sides of most comparisons, and commissions never change verdicts. How we make money
What we know about Fish Audio
Fish Audio is a text-to-speech apis platform: developer-first hosted API from the team behind the open-source fish-speech models; simple prepaid pay-as-you-go billing, 80+ language coverage, and a free fair-use model tier (s2.1-pro-free). Facts here cover the hosted API, not the OSS model. This profile tracks every Fish Audio fact we have verified, each linked to a primary source and dated.
Fish Audio does not publish a public per-minute rate we have been able to verify, so the price figure here stays blank until we can confirm one. We would rather show nothing than a guessed number; if you have a current Fish Audio quote, it is the fastest way for us to close that gap.
On capabilities, Fish Audio covers streaming audio output, realtime websocket api, instant voice cloning, emotion / style controls, and self-host / on-prem option. Each of those is verified against Fish Audio's own docs or dashboard, not marketing copy.
Among the 46 text-to-speech tools in our matrix, Fish Audio leads with the 5th-widest language coverage and lags at the 4th-fastest time-to-first-byte; the fact sheet below has the raw numbers behind that placement.
In total we track 14 verified facts for Fish Audio today, each linking the primary source it came from so you can check our work - and vendor claims we have not measured ourselves are labeled as such on the Fish Audio fact sheet below.
Reviewed by vsref Editorialfacts verified Jul 20, 2026Methodology →
Fish Audio pricing
Usage-priced at $15 per 1M characters (≈ $0.0143 per audio-minute), verified Jul 20, 2026 (source).
| Monthly volume | $ / 1M chars | $ / audio-min | Monthly bill |
|---|---|---|---|
| 200K chars/moHobby project | $15 | $0.0143 | $3 |
| 2M chars/moProduct feature | $15 | $0.0143 | $30 |
| 20M chars/moAt scale | $15 | $0.0142 | $300 |
Effective rates from published pricing; subscription plans resolve to plan fee plus overage. ~950 characters ≈ 1 audio minute. How we compute costs →
Fact sheet
Every row independently verifiedConsidering a switch? Best Fish Audio alternatives →
STT in this stack
Text-to-speech is half of a voice pipeline: the other half is the speech-to-text that listens. Compare transcription engines on accuracy, streaming latency, and per-minute price: Deepgram, AssemblyAI, GPT-4o Transcribe, or the full speech-to-text comparison →
Voice agents in this stack
The engine is one layer: a voice agent speaks through its text-to-speech engine, but orchestration, telephony, and turn-taking come from the agent platform. Compare the platforms builders pair Fish Audio with: Pipecat, OpenAI Realtime API, Twilio ConversationRelay, or the full voice-agent comparison →
Distribute it
Most Fish Audio voiceover ends up in short-form video, and publishing that video across TikTok, YouTube, and Instagram is a scheduling problem with real per-channel pricing. Compare the schedulers creators actually run: Buffer, Postiz, Mixpost, or the social media scheduling platforms compared →
Avatar video in this stack
A cloned or bring-your-own Fish Audio voice does not have to stay audio-only: AI avatar video platforms lip-sync it onto a talking avatar for finished video. Compare the platforms: HeyGen, Synthesia, Hedra, or the full avatar-video comparison →