vsref

12 Best Kokoro Alternatives (2026)

Kokoro is tiny 82m-parameter open-weight tts model that produces high-quality speech fast and cheaply, with 54 preset voices across 8 languages.. Teams that switch usually cite price at production volume, language coverage, self-hosting and license control. The alternatives below are ranked by published head-to-head verdicts, not sponsorship.

Not ready to switch? Full Kokoro review →

If you sign up through links on this page, vsref may earn a commission; programs exist on both sides of most comparisons, and commissions never change verdicts. How we make money

Reviewed by vsref Editorialfacts verified Sep 23, 2026Methodology →

Where to switch, by reason

Switching because of price at production volume →see Cartesia
Switching because of language coverage →see Inworld TTS
Switching because of self-hosting and license control →see Voxtral TTS
Switching because of audiobooks →see Azure Speech
Switching because of content creators →see ElevenLabs
Cloud-utility TTS at commodity pricesvs Kokoro: ~5× the languages, adds streaming audio output.
Enterprise hyperscaler TTS with custom-voice depthvs Kokoro: ~13× the languages, adds streaming audio output.
Localization-first TTS: the MARS 8 family (flash/pro/instruct variants) plus dubbing and translated-TTS pipelines, credit-based plans from $5/mo, aimed at media, sports, and content going multilingual.vs Kokoro: ~18× the languages, adds streaming audio output.
From $5/moTry CAMB.AI →
Lowest-latency TTS for real-time voice agentsvs Kokoro: ~5× the languages, adds streaming audio output.
Production-minded open TTS from a commercial voice company - MIT license, built-in watermarking, and a fast Turbo variant, with Resemble's paid API as the scale-up path.vs Kokoro: ~3× the languages, adds instant voice cloning.
#6ChatTTS logoChatTTSOSS
Optimized for natural dialogue-style speech for LLM assistants; the licensing combination (AGPLv3+ code, CC BY-NC 4.0 weights, research/education only) rules out most commercial SaaS use without a separate deal.vs Kokoro: about half the languages, adds streaming audio output.
Full-stack open TTS from Alibaba's speech team: 9 languages plus 18+ Chinese dialects, 150 ms streaming latency claim, instruction control of emotion/dialect/speed, and training + deployment scripts under Apache-2.0.vs Kokoro: ~15% more languages, adds streaming audio output.
Enterprise real-time voice-agent TTSvs Kokoro: ~15% fewer languages, adds streaming audio output.
Best-known open model for scripted two-speaker dialogue rather than narration - Apache-2.0, English-only, GPU-oriented.vs Kokoro: about half the languages, adds streaming audio output.
Premium AI voice platform for creators and developersvs Kokoro: ~4× the languages, adds streaming audio output.
#11F5-TTS logoF5-TTSOSS
The go-to research-grade voice-cloning model - actively maintained and broadly ported, with the classic code-vs-weights license split: commercial products must retrain or license around the CC-BY-NC checkpoints.vs Kokoro: adds streaming audio output.
See pricingVisit F5-TTS →
Developer-first hosted API from the team behind the open-source fish-speech models; simple prepaid pay-as-you-go billing, 80+ language coverage, and a free fair-use model tier (s2.1-pro-free). Facts here cover the hosted API, not the OSS model.vs Kokoro: ~10× the languages, adds streaming audio output.
All 51 text-to-speech apis alternatives
Top-tier expressive multilingual open-weights TTS whose license moved from permissive to research/non-commercial; commercial use requires a license from Fish Audio or their hosted API.vs Kokoro: ~10× the languages, adds streaming audio output.
Hyperscaler TTS with the broadest voice/language catalogvs Kokoro: ~9× the languages, adds streaming audio output.
The de facto community standard for DIY voice cloning (60k GitHub stars), with a full WebUI covering dataset prep, ASR, training, and inference across five languages; MIT-licensed.vs Kokoro: ~40% fewer languages, adds instant voice cloning.
Real-time voice-agent infrastructure play: WebSocket-first streaming TTS in 5 European languages with instant cloning, on-device models (Phonon), and credit-based pricing. Active and well-funded (site announced funding extension to $100M, July 2026).vs Kokoro: ~40% fewer languages, adds streaming audio output.
Usage-priced hosted TTS from xAI, part of the Grok Voice stack (TTS, STT, and a speech-to-speech Voice Agent API), aimed at developers who want low-latency voice output alongside Grok models.vs Kokoro: ~2.5× the languages, adds streaming audio output.
Strong emotional/expressive open TTS from a well-funded lab - but 'Apache-2.0' only covers the repo code; V2 weights carry a 100k-MAU community license and the newer V3 is research/non-commercial.vs Kokoro: adds streaming audio output.
Expressiveness-first TTS (Octave understands the meaning of the text it speaks); subscription plans gate commercial use, with Octave 2 (preview) adding ~100ms latency and 11 languages for realtime use.vs Kokoro: ~40% more languages, adds streaming audio output.
Emotion-controllable zero-shot voice cloning for production use, but under a custom Bilibili license (not OSI-approved) with scale caps and AI-training restrictions.vs Kokoro: about half the languages, adds instant voice cloning.
Cost-leader realtime TTS for voice agents and games; hybrid pay-as-you-go plus monthly credit plans that lower the per-1M-character rate as commitment grows; SOC 2 Type II with zero-data-retention posture.vs Kokoro: ~25× the languages, adds streaming audio output.
#22KittenTTS logoKittenTTSOSS
The lightweight/edge option: ONNX inference, sub-100 MB downloads, 8 built-in voices - trades voice cloning and multilinguality for footprint.vs Kokoro: about half the languages.
Research-lab open TTS optimized for real-time streaming (Delayed Streams Modeling); Pocket TTS targets on-device/CPU deployment while the larger DSM TTS targets production streaming servers (Rust backend).vs Kokoro: about half the languages, adds streaming audio output.
#24LMNT logoLMNT
Speed-and-affordability challenger: a single fast model (Blizzard 2) with streaming-first APIs, unlimited voice clones on every tier, and flat subscription pricing with per-1K overage.vs Kokoro: ~4× the languages, adds streaming audio output.
From $10/moVisit LMNT →
#25Maya1 logoMaya1OSS
Apache 2.0 expressive English TTS you can run on a single 16GB+ GPU; stands out for natural-language voice design and 20+ inline emotion tags rather than audio-sample cloning, with real-time streaming via vLLM.vs Kokoro: about half the languages, adds streaming audio output.
See pricingVisit Maya1 →
#26MegaTTS3 logoMegaTTS3OSS
Research-grade Apache-2.0 TTS whose practical cloning is gated: the WaveVAE encoder is not released, so users must submit audio to ByteDance channels to obtain pre-extracted speaker latents (.npy) for cloning.vs Kokoro: about half the languages, adds instant voice cloning.
Multilingual cloning-first TTS with aggressive pricingvs Kokoro: ~5× the languages, adds streaming audio output.
API arm of the Murf studio platform: 150+ voices in 35 languages, SSML support, word timestamps, and Falcon 2 aimed at high-concurrency voice agents at $0.01/1K characters. Note: Murf Studio subscription plans (murf.ai/pricing) are a separate product from API pay-as-you-go pricing and API characters are purchased separately.vs Kokoro: ~4× the languages, adds streaming audio output.
Indie batch-content workhorse: aggregates a very broad voice/language catalog for voiceovers, audiobooks and video narration, priced per output minute (prepaid packs, no subscription) rather than per character; not aimed at real-time agent use.vs Kokoro: ~13× the languages, adds streaming audio output.
Hybrid hosted + open-weights play: SSE/WebSocket streaming TTS API at app.neuphonic.com, and tiny CPU-only on-device models (NeuTTS-Air ~360M Apache-2.0, NeuTTS-Nano ~120M) in GGUF for phones/Raspberry Pi - privacy/edge-deployment angle. NOTE: the site's pricing page returned 404 at verification time; hosted-plan pricing treated as not published.vs Kokoro: ~15% fewer languages, adds streaming audio output.
A small, GPU-efficient 9-language TTS checkpoint for teams already in the NVIDIA NeMo/Riva ecosystem; commercially usable open weights, but zero-shot voice cloning was removed from the open release and it caps generations at about 20 seconds.vs Kokoro: ~15% more languages.
Simple usage-based TTS inside a general AI platformvs Kokoro: adds streaming audio output.
The 'LLM-as-TTS' approach under Apache-2.0 - expressive and clonable, but repo activity has slowed since mid-2025.vs Kokoro: adds streaming audio output.
#34OuteTTS logoOuteTTSOSS
The llama.cpp-native option: runs via GGUF on CUDA/ROCm/Vulkan/Metal and even in the browser (Transformers.js). License splits by model: the 0.6B (Qwen3-based) is Apache-2.0, the 1B Llama-based flagship is CC-BY-NC-SA-4.0 (non-commercial).vs Kokoro: ~3× the languages, adds instant voice cloning.
#35Piper logoPiperOSS
The pragmatic embedded/self-hosted choice: no cloning or frills, just quick offline speech in many languages - note the license change from MIT (archived rhasspy/piper) to GPL-3.0 in the successor repo.
See pricingVisit Piper →
#36Qwen3-TTS logoQwen3-TTSOSS
Genuinely open weights (confirmed - not API-only since the Jan 2026 release): Apache-2.0 checkpoints on Hugging Face with an Alibaba Cloud DashScope API for hosted use.vs Kokoro: ~25% more languages, adds streaming audio output.
Security-first enterprise play: generation plus detection/verification in one platform, pay-as-you-go Flex credits, on-prem option, and the MIT-licensed open-source Chatterbox model family.vs Kokoro: adds streaming audio output.
#38Rime logoRime
Enterprise conversational TTS (IVR, contact centers, voice agents) emphasizing ultra-low latency models (Coda, Mist, Arcana) and self-hosted deployment; usage-based pricing with a single published rate.vs Kokoro: ~6× the languages, adds streaming audio output.
See pricingVisit Rime →
The go-to hosted TTS for Indian-language products (Hindi, Tamil, Telugu, Bengali and more) with REST, HTTP streaming and WebSocket APIs and prepaid INR pricing; not aimed at global multilingual coverage.vs Kokoro: ~40% more languages, adds streaming audio output.
Speed- and price-led challenger from an India-focused voice AI startup; Waves is the speech-model API layer (Lightning TTS, Pulse STT), sold pay-as-you-go with $10 free credits, with HIPAA/SOC2/on-prem reserved for the Enterprise plan.vs Kokoro: ~50% more languages, adds streaming audio output.
Developer platform spun out of the Speechify brand: transparent tiered pricing ($10-$499/mo plus per-1M overage), streaming-native Simba 3.2, and a bundled voice-agents product with flat per-minute rates.vs Kokoro: ~4× the languages, adds streaming audio output.
Budget low-latency English TTS for voice agents; priced at a fraction of premium voice APIs, but currently English-only with a small voice set and no SSML or voice cloning.vs Kokoro: about half the languages, adds streaming audio output.
Differentiates on iterative audio *editing* (emotion, style, breathing, laughter, sighs, polyphone pinyin control) rather than plain synthesis; Apache-2.0 with training code (SFT/DPO/GRPO) and vLLM support.vs Kokoro: ~25% fewer languages, adds instant voice cloning.
Expressive character voices (games, content, Korean/Japanese/English markets first, now 31 languages via Supertonic 3) on cheap credit-based subscriptions from $2.99/month; uniquely pairs the hosted API with the open OpenRAIL-M Supertonic model for on-device synthesis.vs Kokoro: ~4× the languages, adds streaming audio output.
Expressiveness play from a Korean AI-actor studio: strongest on emotion prompts/presets and character voices for content and conversational AI; separate consumer studio subscription ($8.99+) and developer API plans (Free/Lite/Plus).vs Kokoro: ~4× the languages, adds streaming audio output.
Pure price play: markets itself as up to 11x cheaper than ElevenLabs with a simple 3-endpoint API (stream/speech/synthesisTasks); smaller voice/language catalog and no voice cloning documented.vs Kokoro: adds streaming audio output.
#47VibeVoice logoVibeVoiceOSS
The open long-form/multi-speaker specialist - MIT weights, but Microsoft pulled the TTS code from the repo in Sept 2025 and frames the models as research-only.vs Kokoro: about half the languages, adds streaming audio output.
A frontier-lab open-weight TTS you can run on a single 16GB GPU; the CC BY-NC license makes it evaluation/research-only, with Mistral's paid API as the commercial route.vs Kokoro: ~15% more languages, adds streaming audio output.
Open-weights-friendly voice cloning TTS from a frontier AI labvs Kokoro: ~15% more languages, adds streaming audio output.
Enterprise voiceover specialist: polished Studio product for L&D/marketing narration with an API on the side; API pricing is contact-sales, and compliance (SOC 2 Type II, GDPR) and ethical voice sourcing are the pitch.vs Kokoro: adds streaming audio output.
The legacy standard for open voice cloning: dormant upstream since Coqui's Jan 2024 shutdown (community fork idiap/coqui-ai-TTS carries maintenance), and CPML weights bar commercial use.vs Kokoro: ~2× the languages, adds instant voice cloning.

How the top Kokoro alternatives compare

The top 6 in depth: where each alternative ranks across the Text-to-speech APIs field we track, which use cases it takes from Kokoro, and what switching gives up.

1. Amazon Polly

Among the 10 text-to-speech tools we track, Amazon Polly has the 1st-cheapest fast-model rate and the 10th-widest language coverage - a fit for multilingual and localization projects.

Seen from the other side, Kokoro vs Amazon Polly: about half the languages, no streaming audio output.

Amazon Polly lists $30 per 1M characters, verified July 2026.

Amazon Polly cost at monthly volume tiers
Monthly volume$ / 1M chars$ / audio-minMonthly bill
200K chars/moHobby project$30$0.0285$6
2M chars/moProduct feature$30$0.0285$60
20M chars/moAt scale$30$0.0285$600

Usage-priced at $30 per 1M characters (≈ $0.0285 per audio-minute), verified Jul 20, 2026. Effective rates from published pricing; subscription plans resolve to plan fee plus overage. ~950 characters ≈ 1 audio minute. How we compute costs →

Amazon Polly review →

2. Azure Speech

Among the 46 text-to-speech tools we track, Azure Speech has the 3rd-widest language coverage and the 6th-cheapest flagship rate - a fit for multilingual and localization projects and cost-sensitive, high-volume work.

The reverse angle matters too - Kokoro vs Azure Speech: about half the languages, no streaming audio output.

Published pricing starts at $22 per 1M characters, verified July 2026.

Azure Speech cost at monthly volume tiers
Monthly volume$ / 1M chars$ / audio-minMonthly bill
200K chars/moHobby project$22$0.0209$4.40
2M chars/moProduct feature$22$0.0209$44
20M chars/moAt scale$22$0.0209$440

Usage-priced at $22 per 1M characters (≈ $0.0209 per audio-minute), verified Jul 20, 2026. Effective rates from published pricing; subscription plans resolve to plan fee plus overage. ~950 characters ≈ 1 audio minute. How we compute costs →

Azure Speech review →

3. CAMB.AI

Among the 46 text-to-speech tools we track, CAMB.AI has the 2nd-widest language coverage - a fit for multilingual and localization projects.

Before switching, weigh what stays behind - Kokoro vs CAMB.AI: about half the languages, no streaming audio output.

CAMB.AI cost at monthly volume tiers
Monthly volume$ / 1M chars$ / audio-minMonthly bill
200K chars/moHobby project$25$0.0238$5
2M chars/moProduct feature$2.50$0.0024$5
20M chars/moAt scale$0.25$0.0002$5

Cheapest paid plan $5/mo, verified Jul 20, 2026. Effective rates from published pricing; subscription plans resolve to plan fee plus overage. ~950 characters ≈ 1 audio minute. How we compute costs →

CAMB.AI review →

4. Cartesia

Among the 46 text-to-speech tools we track, Cartesia has the 9th-widest language coverage and the 3rd-fastest time-to-first-byte - a fit for multilingual and localization projects and real-time, conversational apps.

Seen from the other side, Kokoro vs Cartesia: about half the languages, no streaming audio output.

Cartesia review →

5. Chatterbox

Among the 46 text-to-speech tools we track, Chatterbox has the 18th-widest language coverage.

Seen from the other side, Kokoro vs Chatterbox: about half the languages, no instant voice cloning.

Chatterbox review →

6. ChatTTS

Among the 46 text-to-speech tools we track, ChatTTS has the 38th-widest language coverage.

The reverse angle matters too - Kokoro vs ChatTTS: ~4× the languages, no streaming audio output.

ChatTTS review →