vsref
Gradium Speech-to-Text logo

Gradium Speech-to-Text Review

Low-latency STT for voice agents

Among the 43 speech-to-text tools we track, Gradium Speech-to-Text has the 40th-widest language coverage.

From $13/mo

Facts verified Jul 20, 2026Try Gradium Speech-to-Text →

If you sign up through links on this page, vsref may earn a commission; programs exist on both sides of most comparisons, and commissions never change verdicts. How we make money

What we know about Gradium Speech-to-Text

This is our verified profile of Gradium Speech-to-Text, a speech-to-text apis platform - low-latency STT for voice agents. Every fact about Gradium Speech-to-Text below carries the source it came from and the day we checked it.

On pricing, Gradium Speech-to-Text starts at $13 per month for its entry tier. That is the sticker rate: real production cost usually runs higher once you add a language model, a voice provider, and telephony minutes.

On capabilities, Gradium Speech-to-Text covers speech translation and websocket streaming api, and does not offer language auto-detection, word-level timestamps, and custom vocabulary / keyterm boosting. Each of those is verified against Gradium Speech-to-Text's own docs or dashboard, not marketing copy.

Placed against the 43 speech-to-text tools we track, Gradium Speech-to-Text's strongest showing is the 40th-widest language coverage - a spread worth weighing against your own priorities.

In total we track 23 verified facts for Gradium Speech-to-Text today, each linking the primary source it came from so you can check our work - and vendor claims we have not measured ourselves are labeled as such on the Gradium Speech-to-Text fact sheet below.

Reviewed by vsref Editorialfacts verified Jul 20, 2026Methodology →

Fact sheet

Pricing
Pricing facts
Pricing modelhybridJul 20
Cheapest paid plan$13Jul 20
Free tier quota45k credits (~4 hrs STT)Jul 20
Published volume discountsYes - add-on credits cheaper at higher tiersJul 20
Concurrency on base planFree 3; XS/S 20 sessionsJul 20
Capabilities
Capabilities facts
WER (third-party benchmark)8.38 % WERJul 20
Streaming latency (vendor-claimed)~300 msJul 20
Languages supported5Jul 20
Language auto-detection✗ NoJul 20
Speaker diarization✗ Not availableJul 20
PII redaction✗ Not availableJul 20
Word-level timestamps✗ NoJul 20
Custom vocabulary / keyterm boosting✗ NoJul 20
Entity detection✗ Not offered.Jul 20
Sentiment analysis✗ Not offered.Jul 20
Summarization endpoint✗ Not offered.Jul 20
Speech translation✓ YesJul 20
Priced audio-intelligence add-onsNone (semantic VAD, translation only)Jul 20
Compliance & trust
Compliance & trust facts
Self-host / on-prem option✗ NoJul 20
Build experience
Build experience facts
Official SDKsPython (pip install gradium)Jul 20
Max file size / duration300 s per session (5 min)Jul 20
Supported audio formatsWAV, PCM, Opus, mu-law, A-lawJul 20
Websocket streaming API✓ YesJul 20

Considering a switch? Best Gradium Speech-to-Text alternatives →

Voice agents in this stack

The engine is one layer: a voice agent hears through its transcription engine, but orchestration, telephony, and turn-taking come from the agent platform. Compare the platforms builders pair Gradium Speech-to-Text with: Pipecat, OpenAI Realtime API, Twilio ConversationRelay, or the full voice-agent comparison →

Want it done for you?

Gradium Speech-to-Text gets you a transcript; the bot that joins the call, labels speakers, and writes the summary is still your build. If that is more pipeline than you want to own, AI meeting notetakers do the whole job end to end: Otter.ai, Fireflies.ai, Fathom, or the full notetaker comparison.