Cartesia Review
Lowest-latency TTS for real-time voice agents
Among the 46 text-to-speech tools we track, Cartesia has the 9th-widest language coverage and the 3rd-fastest time-to-first-byte - a fit for multilingual and localization projects and real-time, conversational apps.
From $5/mo
If you sign up through links on this page, vsref may earn a commission; programs exist on both sides of most comparisons, and commissions never change verdicts. How we make money
What we know about Cartesia
This is our verified profile of Cartesia, a text-to-speech apis platform - lowest-latency TTS for real-time voice agents. Every fact about Cartesia below carries the source it came from and the day we checked it.
On pricing, Cartesia starts at $5 per month for its entry tier. That is the sticker rate: real production cost usually runs higher once you add a language model, a voice provider, and telephony minutes.
On capabilities, Cartesia covers streaming audio output, realtime websocket api, instant voice cloning, professional voice cloning, word-level timestamps, and pronunciation dictionaries, and does not offer commercial use on free tier. Each of those is verified against Cartesia's own docs or dashboard, not marketing copy.
For compliance, with Cartesia: SOC 2 Type II is in place. If you are in a regulated space, confirm the current posture with Cartesia before you commit, since these change plan by plan.
Placed against the 46 text-to-speech tools we track, Cartesia's strongest showing is the 9th-widest language coverage, while it trails at 3rd on latency - a spread worth weighing against your own priorities.
In total we track 21 verified facts for Cartesia today, each linking the primary source it came from so you can check our work - and vendor claims we have not measured ourselves are labeled as such on the Cartesia fact sheet below.
Reviewed by vsref Editorialfacts verified Jul 20, 2026Methodology →
Cartesia pricing
Cheapest paid plan $5/mo with 100,000 characters included, verified Jul 20, 2026 (source).
| Monthly volume | $ / 1M chars | $ / audio-min | Monthly bill |
|---|---|---|---|
| 200K chars/moHobby project | Higher plan | Higher plan | Higher plan |
| 2M chars/moProduct feature | Higher plan | Higher plan | Higher plan |
| 20M chars/moAt scale | Higher plan | Higher plan | Higher plan |
Effective rates from published pricing; subscription plans resolve to plan fee plus overage. ~950 characters ≈ 1 audio minute. How we compute costs →
Fact sheet
Every row independently verifiedConsidering a switch? Best Cartesia alternatives →
STT in this stack
Text-to-speech is half of a voice pipeline: the other half is the speech-to-text that listens. Compare transcription engines on accuracy, streaming latency, and per-minute price: Deepgram, AssemblyAI, GPT-4o Transcribe, or the full speech-to-text comparison →
Voice agents in this stack
The engine is one layer: a voice agent speaks through its text-to-speech engine, but orchestration, telephony, and turn-taking come from the agent platform. Compare the platforms builders pair Cartesia with: Pipecat, OpenAI Realtime API, Twilio ConversationRelay, or the full voice-agent comparison →
Distribute it
Most Cartesia voiceover ends up in short-form video, and publishing that video across TikTok, YouTube, and Instagram is a scheduling problem with real per-channel pricing. Compare the schedulers creators actually run: Buffer, Postiz, Mixpost, or the social media scheduling platforms compared →
Avatar video in this stack
A cloned or bring-your-own Cartesia voice does not have to stay audio-only: AI avatar video platforms lip-sync it onto a talking avatar for finished video. Compare the platforms: HeyGen, Synthesia, Hedra, or the full avatar-video comparison →