vsref
Gradium logo

Gradium Review

Real-time voice-agent infrastructure play: WebSocket-first streaming TTS in 5 European languages with instant cloning, on-device models (Phonon), and credit-based pricing. Active and well-funded (site announced funding extension to $100M, July 2026).

Among the 15 text-to-speech tools we track, Gradium has the 10th-largest voice library.

From $13/mo

Facts verified Jul 20, 2026Try Gradium →

If you sign up through links on this page, vsref may earn a commission; programs exist on both sides of most comparisons, and commissions never change verdicts. How we make money

What we know about Gradium

This is our verified profile of Gradium, a text-to-speech apis platform - real-time voice-agent infrastructure play: WebSocket-first streaming TTS in 5 European languages with instant cloning, on-device models (Phonon), and credit-based pricing. Active and well-funded (site announced funding extension to $100M, July 2026). Every fact about Gradium below carries the source it came from and the day we checked it.

On pricing, Gradium starts at $13 per month for its entry tier. That is the sticker rate: real production cost usually runs higher once you add a language model, a voice provider, and telephony minutes.

On capabilities, Gradium covers streaming audio output, realtime websocket api, instant voice cloning, professional voice cloning, ssml support, and word-level timestamps, and does not offer commercial use on free tier. Each of those is verified against Gradium's own docs or dashboard, not marketing copy.

Among the 15 text-to-speech tools in our matrix, Gradium leads with the 10th-largest voice library and lags at the 36th-widest language coverage; the fact sheet below has the raw numbers behind that placement.

In total we track 21 verified facts for Gradium today, each linking the primary source it came from so you can check our work - and vendor claims we have not measured ourselves are labeled as such on the Gradium fact sheet below.

Reviewed by vsref Editorialfacts verified Jul 20, 2026Methodology →

Gradium pricing

Cheapest paid plan $13/mo with 225,000 characters included, verified Jul 20, 2026 (source).

Gradium cost at monthly volume tiers
Monthly volume$ / 1M chars$ / audio-minMonthly bill
200K chars/moHobby project$65$0.0618$13
2M chars/moProduct featureHigher planHigher planHigher plan
20M chars/moAt scaleHigher planHigher planHigher plan

Effective rates from published pricing; subscription plans resolve to plan fee plus overage. ~950 characters ≈ 1 audio minute. How we compute costs →

Fact sheet

Pricing
Pricing facts
Pricing modelMonthly subscription credit plans: Free $0 (45k credits)Jul 20
Cheapest paid plan$13Jul 20
Characters included on cheapest plan225,000 chars/moJul 20
Free tier quota45,000 credits (~1 hour of audio)Jul 20
Enterprise / contact-sales thresholdEnterprise plan: custom pricingJul 20
Capabilities
Capabilities facts
Streaming audio output✓ YesJul 20
Realtime websocket API✓ YesJul 20
Voice library size66 flagship voicesJul 20
Instant voice cloning✓ YesJul 20
Professional voice cloning✓ YesJul 20
Minimum audio for voice cloning10 secondsJul 20
Languages supported5 languagesJul 20
SSML support✓ partialJul 20
Word-level timestamps✓ YesJul 20
Pronunciation dictionaries✓ YesJul 20
Compliance & trust
Compliance & trust facts
Commercial use on free tier✗ NoJul 20
Self-host / on-prem option✓ YesJul 20
Build experience
Build experience facts
Official SDKsPython SDK (documented)Jul 20
Concurrency on base plan2-5 simultaneous TTS requests (Free through S); M: 10; L: 15Jul 20
Output formatspcm (48kHz defaultJul 20
Max input per requestSingle TTS/STT session up to 300 secondsJul 20

Considering a switch? Best Gradium alternatives →

STT in this stack

Text-to-speech is half of a voice pipeline: the other half is the speech-to-text that listens. Compare transcription engines on accuracy, streaming latency, and per-minute price: Deepgram, AssemblyAI, GPT-4o Transcribe, or the full speech-to-text comparison →

Voice agents in this stack

The engine is one layer: a voice agent speaks through its text-to-speech engine, but orchestration, telephony, and turn-taking come from the agent platform. Compare the platforms builders pair Gradium with: Pipecat, OpenAI Realtime API, Twilio ConversationRelay, or the full voice-agent comparison →

Distribute it

Most Gradium voiceover ends up in short-form video, and publishing that video across TikTok, YouTube, and Instagram is a scheduling problem with real per-channel pricing. Compare the schedulers creators actually run: Buffer, Postiz, Mixpost, or the social media scheduling platforms compared →

Avatar video in this stack

A cloned or bring-your-own Gradium voice does not have to stay audio-only: AI avatar video platforms lip-sync it onto a talking avatar for finished video. Compare the platforms: HeyGen, Synthesia, Hedra, or the full avatar-video comparison →