vsref

Best Speech-to-text APIs for Voice Agents (2026)

For voice agents, Cartesia Ink is our pick (from $5/mo): For voice agents, streaming latency is the top priority. Live transcription feeding a voice bot, where streaming latency decides whether the agent interrupts or lags the caller. Below is the full ranking and the tradeoffs, or read how we score.

If you sign up through links on this page, vsref may earn a commission; programs exist on both sides of most comparisons, and commissions never change verdicts. How we make money

Reviewed by vsref Editorialfacts verified Jul 24, 2026Methodology →

Streaming STT for voice agents with native turn detection2 of 2 points · 1 matchup
Flagship GPT-4o based transcription API from OpenAI2 of 2 points · 1 matchup
Open-weights multilingual ASR models2 of 2 points · 1 matchup
See pricingWebsite →

What matters for voice agents

Weighted attribute comparison for Voice Agents
FactCartesia InkOpenAI gpt-4o-transcribeQwen3-ASRAssemblyAIDeepgram
Streaming latency (vendor-claimed)×5~88 msJul 20n/an/a~150 msJul 20~300 msJul 20
Websocket streaming API×5✓ YesJul 20✓ YesJul 20✓ YesJul 20✓ YesJul 20✓ YesJul 20
Streaming price per audio minute×4n/an/a0.005 $/audio-minJul 200.008 $/audio-minJul 200.005 $/audio-minJul 20
Concurrency on base plan×3Free: 8, Pro: 12 concurrent STT requestsJul 20Tier-based rate limits; Tier 1: 500 RPM / 10,000 TPMJul 20n/aAsync: 200+ concurrent jobs; streaming: 100 new streams/minJul 20PAYG STT: 50 REST, 150 websocket concurrentJul 20
Custom vocabulary / keyterm boosting×2✓ YesJul 20✓ YesJul 20n/a✓ YesJul 20✓ YesJul 20
Swipe → to see every tool column.
×5 Streaming latency (vendor-claimed): The transcript delay sits inside every conversational turn; it is the floor on how fast the agent can respond.×5 Websocket streaming API: A persistent streaming connection is the integration model voice agents are built on; batch-only APIs are a hard stop.×4 Streaming price per audio minute: Streaming rates run higher than batch and accrue on every call minute.×3 Concurrency on base plan: Concurrent-stream caps decide how many simultaneous calls you can serve before a bigger contract.×2 Custom vocabulary / keyterm boosting: Boosting product names and jargon cuts the misrecognitions that derail an automated call.

The ranking, tool by tool

For voice agents, streaming latency is the top priority.

For voice agents, streaming latency is the top priority. Full Cartesia Ink vs Deepgram verdict →

For live voice agent use cases, real-time streaming capability is the decisive factor.

For live voice agent use cases, real-time streaming capability is the decisive factor. Full OpenAI gpt-4o-transcribe vs OpenAI Whisper (API) verdict →

For voice agents, a native WebSocket streaming API is essential for real-time transcription.
See pricingWebsite →

For voice agents, a native WebSocket streaming API is essential for real-time transcription. Full Qwen3-ASR vs OpenAI Whisper (API) verdict →

For voice agents, streaming capability is the decisive factor.

For voice agents, streaming capability is the decisive factor. Full AssemblyAI vs OpenAI Whisper (API) verdict →

Both tools offer WebSocket streaming APIs, so that factor cancels out. Full AssemblyAI vs Speechmatics verdict →

For voice agents, streaming latency is the top-weighted factor. Full AssemblyAI vs Soniox verdict →

For voice agents, streaming latency is the top priority. Full AssemblyAI vs Rev AI verdict →

For voice agents, streaming latency is the decisive attribute. Full AssemblyAI vs Gladia verdict →

For voice agents, streaming latency is the top-weighted factor.
See pricingTry Deepgram

For voice agents, streaming latency is the top-weighted factor. Full Deepgram vs Amazon Transcribe verdict →

For voice agents, streaming latency is the top-weighted factor. Full Deepgram vs Azure AI Speech (STT) verdict →

For voice agents, streaming latency is the top factor. Full Deepgram vs Rev AI verdict →

For voice agents, streaming latency and websocket support are the two heaviest factors, and Deepgram dominates both. Full Deepgram vs NVIDIA Parakeet / Riva verdict →

Both tools offer websocket streaming and custom vocabulary, so those attributes are tied. Full Deepgram vs xAI Grok Speech-to-Text verdict →

For voice agents, streaming capability and latency are the top two factors, and Deepgram leads decisively on both. Full Deepgram vs OpenAI Whisper (API) verdict →

For voice agents, streaming latency is the decisive factor. Full Deepgram vs Speechmatics verdict →

For voice agents, streaming latency and websocket support are the heaviest factors. Full Deepgram vs Google Cloud Speech-to-Text verdict →

Both tools match on streaming latency at 300 ms and both offer websocket streaming APIs, so the top two attributes are tied. Full Deepgram vs Gladia verdict →

For voice agents, streaming latency and websocket support are the two most critical attributes, both weighted 5 out of 5. Full Deepgram vs Cohere Transcribe verdict →

On the two heaviest attributes, both tools tie: streaming latency is vendor-claimed at 150 ms each, and both offer WebSocket streaming APIs.

On the two heaviest attributes, both tools tie: streaming latency is vendor-claimed at 150 ms each, and both offer WebSocket streaming APIs. Full ElevenLabs Scribe vs AssemblyAI verdict →

For voice agents, streaming latency and WebSocket support are the two most critical attributes. Full ElevenLabs Scribe vs OpenAI Whisper (API) verdict →

On the two heaviest attributes, ElevenLabs Scribe leads on streaming latency at 150 ms versus Mistral Voxtral Transcribe at 200 ms, and both tools offer a WebSocket streaming API. Full ElevenLabs Scribe vs Mistral Voxtral Transcribe verdict →

For voice agents, streaming latency and WebSocket availability carry the most weight.

For voice agents, streaming latency and WebSocket availability carry the most weight. Full Mistral Voxtral Transcribe vs Deepgram verdict →

For live voice agents, streaming capability and latency are decisive. Full Mistral Voxtral Transcribe vs OpenAI Whisper (API) verdict →

For voice agents, a WebSocket streaming API is essential for low-latency bidirectional communication.

For voice agents, a WebSocket streaming API is essential for low-latency bidirectional communication. Full Amazon Transcribe vs Google Cloud Speech-to-Text verdict →

For live voice agents, a native WebSocket streaming API is critical because it enables true bidirectional, low-latency communication.

For live voice agents, a native WebSocket streaming API is critical because it enables true bidirectional, low-latency communication. Full Azure AI Speech (STT) vs Google Cloud Speech-to-Text verdict →

For voice agents, streaming latency and a WebSocket API are the two most critical requirements.
See pricingTry Gladia

For voice agents, streaming latency and a WebSocket API are the two most critical requirements. Full Gladia vs OpenAI Whisper (API) verdict →

On the heaviest attribute, streaming latency, Soniox claims 249 ms versus Deepgram's 300 ms, a meaningful gap when a voice agent must decide whether to interrupt a caller.
See pricingTry Soniox

On the heaviest attribute, streaming latency, Soniox claims 249 ms versus Deepgram's 300 ms, a meaningful gap when a voice agent must decide whether to interrupt a caller. Full Soniox vs Deepgram verdict →

Ultra-fast, low-cost hosted Whisper transcription API (no realtime streaming).

Ultra-fast, low-cost hosted Whisper transcription API (no realtime streaming). No won verdicts for this use case yet; it ranks on ties and near-misses.

On-device streaming STT for live voice interfaces, from tiny edge models to Whisper Large v3-beating accuracy.
See pricingWebsite →

On-device streaming STT for live voice interfaces, from tiny edge models to Whisper Large v3-beating accuracy. No won verdicts for this use case yet; it ranks on ties and near-misses.

Open-weights, GPU-accelerated self-hosted STT stack.

Open-weights, GPU-accelerated self-hosted STT stack. No won verdicts for this use case yet; it ranks on ties and near-misses.

Low-cost pay-as-you-go file transcription from OpenAI; no first-party streaming or diarization.

Low-cost pay-as-you-go file transcription from OpenAI; no first-party streaming or diarization. No won verdicts for this use case yet; it ranks on ties and near-misses.

Enterprise-grade open ASR for accurate batch transcription.

Enterprise-grade open ASR for accurate batch transcription. No won verdicts for this use case yet; it ranks on ties and near-misses.

Hyperscaler STT API with Chirp foundation models and enterprise compliance.

Hyperscaler STT API with Chirp foundation models and enterprise compliance. No won verdicts for this use case yet; it ranks on ties and near-misses.

Low-cost hosted STT API on the Grok stack.

Low-cost hosted STT API on the Grok stack. No won verdicts for this use case yet; it ranks on ties and near-misses.

Transcription-heritage STT API with low per-hour pricing and open (non-commercial) Reverb models.
See pricingTry Rev AI

Transcription-heritage STT API with low per-hour pricing and open (non-commercial) Reverb models. No won verdicts for this use case yet; it ranks on ties and near-misses.

Accuracy-first enterprise STT with flexible deployment (SaaS, container, on-prem).

Accuracy-first enterprise STT with flexible deployment (SaaS, container, on-prem). No won verdicts for this use case yet; it ranks on ties and near-misses.

More Speech-to-text APIs buyer guides