Best Speech-to-text APIs for Voice Agents (2026)
For voice agents, Cartesia Ink is our pick (from $5/mo): For voice agents, streaming latency is the top priority. Live transcription feeding a voice bot, where streaming latency decides whether the agent interrupts or lags the caller. Below is the full ranking and the tradeoffs, or read how we score.
If you sign up through links on this page, vsref may earn a commission; programs exist on both sides of most comparisons, and commissions never change verdicts. How we make money
Reviewed by vsref Editorialfacts verified Jul 24, 2026Methodology →
What matters for voice agents
Weight ×5 = decisive, ×1 = relevant| Fact | Cartesia Ink | OpenAI gpt-4o-transcribe | Qwen3-ASR | AssemblyAI | Deepgram |
|---|---|---|---|---|---|
| Streaming latency (vendor-claimed)×5 | ~88 msJul 20 | n/a | n/a | ~150 msJul 20 | ~300 msJul 20 |
| Websocket streaming API×5 | ✓ YesJul 20 | ✓ YesJul 20 | ✓ YesJul 20 | ✓ YesJul 20 | ✓ YesJul 20 |
| Streaming price per audio minute×4 | n/a | n/a | 0.005 $/audio-minJul 20 | 0.008 $/audio-minJul 20 | 0.005 $/audio-minJul 20 |
| Concurrency on base plan×3 | Free: 8, Pro: 12 concurrent STT requestsJul 20 | Tier-based rate limits; Tier 1: 500 RPM / 10,000 TPMJul 20 | n/a | Async: 200+ concurrent jobs; streaming: 100 new streams/minJul 20 | PAYG STT: 50 REST, 150 websocket concurrentJul 20 |
| Custom vocabulary / keyterm boosting×2 | ✓ YesJul 20 | ✓ YesJul 20 | n/a | ✓ YesJul 20 | ✓ YesJul 20 |
The ranking, tool by tool
For voice agents, streaming latency is the top priority. Full Cartesia Ink vs Deepgram verdict →
For live voice agent use cases, real-time streaming capability is the decisive factor. Full OpenAI gpt-4o-transcribe vs OpenAI Whisper (API) verdict →
For voice agents, a native WebSocket streaming API is essential for real-time transcription. Full Qwen3-ASR vs OpenAI Whisper (API) verdict →
For voice agents, streaming capability is the decisive factor. Full AssemblyAI vs OpenAI Whisper (API) verdict →
Both tools offer WebSocket streaming APIs, so that factor cancels out. Full AssemblyAI vs Speechmatics verdict →
For voice agents, streaming latency is the top-weighted factor. Full AssemblyAI vs Soniox verdict →
For voice agents, streaming latency is the top priority. Full AssemblyAI vs Rev AI verdict →
For voice agents, streaming latency is the decisive attribute. Full AssemblyAI vs Gladia verdict →
For voice agents, streaming latency is the top-weighted factor. Full Deepgram vs Amazon Transcribe verdict →
For voice agents, streaming latency is the top-weighted factor. Full Deepgram vs Azure AI Speech (STT) verdict →
For voice agents, streaming latency is the top factor. Full Deepgram vs Rev AI verdict →
For voice agents, streaming latency and websocket support are the two heaviest factors, and Deepgram dominates both. Full Deepgram vs NVIDIA Parakeet / Riva verdict →
Both tools offer websocket streaming and custom vocabulary, so those attributes are tied. Full Deepgram vs xAI Grok Speech-to-Text verdict →
For voice agents, streaming capability and latency are the top two factors, and Deepgram leads decisively on both. Full Deepgram vs OpenAI Whisper (API) verdict →
For voice agents, streaming latency is the decisive factor. Full Deepgram vs Speechmatics verdict →
For voice agents, streaming latency and websocket support are the heaviest factors. Full Deepgram vs Google Cloud Speech-to-Text verdict →
Both tools match on streaming latency at 300 ms and both offer websocket streaming APIs, so the top two attributes are tied. Full Deepgram vs Gladia verdict →
For voice agents, streaming latency and websocket support are the two most critical attributes, both weighted 5 out of 5. Full Deepgram vs Cohere Transcribe verdict →
On the two heaviest attributes, both tools tie: streaming latency is vendor-claimed at 150 ms each, and both offer WebSocket streaming APIs. Full ElevenLabs Scribe vs AssemblyAI verdict →
For voice agents, streaming latency and WebSocket support are the two most critical attributes. Full ElevenLabs Scribe vs OpenAI Whisper (API) verdict →
On the two heaviest attributes, ElevenLabs Scribe leads on streaming latency at 150 ms versus Mistral Voxtral Transcribe at 200 ms, and both tools offer a WebSocket streaming API. Full ElevenLabs Scribe vs Mistral Voxtral Transcribe verdict →
For voice agents, streaming latency and WebSocket availability carry the most weight. Full Mistral Voxtral Transcribe vs Deepgram verdict →
For live voice agents, streaming capability and latency are decisive. Full Mistral Voxtral Transcribe vs OpenAI Whisper (API) verdict →
For voice agents, a WebSocket streaming API is essential for low-latency bidirectional communication. Full Amazon Transcribe vs Google Cloud Speech-to-Text verdict →
For live voice agents, a native WebSocket streaming API is critical because it enables true bidirectional, low-latency communication. Full Azure AI Speech (STT) vs Google Cloud Speech-to-Text verdict →
For voice agents, streaming latency and a WebSocket API are the two most critical requirements. Full Gladia vs OpenAI Whisper (API) verdict →
On the heaviest attribute, streaming latency, Soniox claims 249 ms versus Deepgram's 300 ms, a meaningful gap when a voice agent must decide whether to interrupt a caller. Full Soniox vs Deepgram verdict →
Ultra-fast, low-cost hosted Whisper transcription API (no realtime streaming). No won verdicts for this use case yet; it ranks on ties and near-misses.
On-device streaming STT for live voice interfaces, from tiny edge models to Whisper Large v3-beating accuracy. No won verdicts for this use case yet; it ranks on ties and near-misses.
Open-weights, GPU-accelerated self-hosted STT stack. No won verdicts for this use case yet; it ranks on ties and near-misses.
Low-cost pay-as-you-go file transcription from OpenAI; no first-party streaming or diarization. No won verdicts for this use case yet; it ranks on ties and near-misses.
Enterprise-grade open ASR for accurate batch transcription. No won verdicts for this use case yet; it ranks on ties and near-misses.
Hyperscaler STT API with Chirp foundation models and enterprise compliance. No won verdicts for this use case yet; it ranks on ties and near-misses.
Low-cost hosted STT API on the Grok stack. No won verdicts for this use case yet; it ranks on ties and near-misses.
Transcription-heritage STT API with low per-hour pricing and open (non-commercial) Reverb models. No won verdicts for this use case yet; it ranks on ties and near-misses.
Accuracy-first enterprise STT with flexible deployment (SaaS, container, on-prem). No won verdicts for this use case yet; it ranks on ties and near-misses.