vsref
Cartesia logo

Cartesia Review

Lowest-latency TTS for real-time voice agents

Among the 46 text-to-speech tools we track, Cartesia has the 9th-widest language coverage and the 3rd-fastest time-to-first-byte - a fit for multilingual and localization projects and real-time, conversational apps.

From $5/mo

Facts verified Jul 20, 2026Try Cartesia →

If you sign up through links on this page, vsref may earn a commission; programs exist on both sides of most comparisons, and commissions never change verdicts. How we make money

What we know about Cartesia

This is our verified profile of Cartesia, a text-to-speech apis platform - lowest-latency TTS for real-time voice agents. Every fact about Cartesia below carries the source it came from and the day we checked it.

On pricing, Cartesia starts at $5 per month for its entry tier. That is the sticker rate: real production cost usually runs higher once you add a language model, a voice provider, and telephony minutes.

On capabilities, Cartesia covers streaming audio output, realtime websocket api, instant voice cloning, professional voice cloning, word-level timestamps, and pronunciation dictionaries, and does not offer commercial use on free tier. Each of those is verified against Cartesia's own docs or dashboard, not marketing copy.

For compliance, with Cartesia: SOC 2 Type II is in place. If you are in a regulated space, confirm the current posture with Cartesia before you commit, since these change plan by plan.

Placed against the 46 text-to-speech tools we track, Cartesia's strongest showing is the 9th-widest language coverage, while it trails at 3rd on latency - a spread worth weighing against your own priorities.

In total we track 21 verified facts for Cartesia today, each linking the primary source it came from so you can check our work - and vendor claims we have not measured ourselves are labeled as such on the Cartesia fact sheet below.

Reviewed by vsref Editorialfacts verified Jul 20, 2026Methodology →

Cartesia pricing

Cheapest paid plan $5/mo with 100,000 characters included, verified Jul 20, 2026 (source).

Cartesia cost at monthly volume tiers
Monthly volume$ / 1M chars$ / audio-minMonthly bill
200K chars/moHobby projectHigher planHigher planHigher plan
2M chars/moProduct featureHigher planHigher planHigher plan
20M chars/moAt scaleHigher planHigher planHigher plan

Effective rates from published pricing; subscription plans resolve to plan fee plus overage. ~950 characters ≈ 1 audio minute. How we compute costs →

Fact sheet

Pricing
Pricing facts
Pricing modelcreditsJul 20
Cheapest paid plan$5Jul 20
Characters included on cheapest plan100,000 chars/moJul 20
Free tier quota20,000 credits/month (~27 TTS minutes)Jul 20
Capabilities
Capabilities facts
TTFB latency (vendor-claimed)~90 msJul 20
Streaming audio output✓ YesJul 20
Realtime websocket API✓ YesJul 20
Instant voice cloning✓ YesJul 20
Professional voice cloning✓ YesJul 20
Minimum audio for voice cloningIVC: clip up to 10 secondsJul 20
Languages supported42 languagesJul 20
Word-level timestamps✓ YesJul 20
Pronunciation dictionaries✓ YesJul 20
Compliance & trust
Compliance & trust facts
Commercial use on free tier✗ NoJul 20
SOC 2 Type II✓ YesJul 20
HIPAA BAA available◑ Enterprise onlyJul 20
Self-host / on-prem option✓ YesJul 20
Model weights licenseclosedJul 20
Build experience
Build experience facts
Official SDKsPython, JavaScript/TypeScriptJul 20
Concurrency on base planFree: 2Jul 20
Output formatsraw PCM (pcm_f32leJul 20

Considering a switch? Best Cartesia alternatives →

STT in this stack

Text-to-speech is half of a voice pipeline: the other half is the speech-to-text that listens. Compare transcription engines on accuracy, streaming latency, and per-minute price: Deepgram, AssemblyAI, GPT-4o Transcribe, or the full speech-to-text comparison →

Voice agents in this stack

The engine is one layer: a voice agent speaks through its text-to-speech engine, but orchestration, telephony, and turn-taking come from the agent platform. Compare the platforms builders pair Cartesia with: Pipecat, OpenAI Realtime API, Twilio ConversationRelay, or the full voice-agent comparison →

Distribute it

Most Cartesia voiceover ends up in short-form video, and publishing that video across TikTok, YouTube, and Instagram is a scheduling problem with real per-channel pricing. Compare the schedulers creators actually run: Buffer, Postiz, Mixpost, or the social media scheduling platforms compared →

Avatar video in this stack

A cloned or bring-your-own Cartesia voice does not have to stay audio-only: AI avatar video platforms lip-sync it onto a talking avatar for finished video. Compare the platforms: HeyGen, Synthesia, Hedra, or the full avatar-video comparison →

Cartesia head-to-head

Cartesia vs ElevenLabs →won 3 · lost 3 · tied 0
AudiobookslostContent CreatorslostDeveloperslostDubbingwonSelf-HostedwonVoice Agentswon
Cartesia vs OpenAI TTS →won 4 · lost 1 · tied 0
AudiobookslostContent CreatorswonDubbingwonSelf-HostedwonVoice Agentswon
Cartesia vs Rime →won 2 · lost 2 · tied 1
AudiobookslostDeveloperswonDubbinglostSelf-HostedtieVoice Agentswon
Cartesia vs Deepgram Aura-2 →won 3 · lost 0 · tied 1
Content CreatorswonDubbingwonSelf-HostedtieVoice Agentswon
Cartesia vs Voxtral TTS →won 5 · lost 1 · tied 0
AudiobookswonContent CreatorswonDeveloperswonDubbingwonSelf-HostedlostVoice Agentswon
Cartesia vs Inworld TTS →won 3 · lost 3 · tied 0
AudiobookslostContent CreatorslostDeveloperswonDubbinglostSelf-HostedwonVoice Agentswon
Cartesia vs LMNT →won 5 · lost 1 · tied 0
AudiobookslostContent CreatorswonDeveloperswonDubbingwonSelf-HostedwonVoice Agentswon