# Best Speech-to-text APIs for Voice Agents (2026)

> The best Speech-to-text APIs platforms for voice agents: Cartesia Ink leads, for live transcription feeding a voice bot, where streaming latency decides.

For voice agents, **Cartesia Ink** is our pick (from $5/mo): For voice agents, streaming latency is the top priority. Live transcription feeding a voice bot, where streaming latency decides whether the agent interrupts or lags the caller. Below is the full ranking and the tradeoffs, or read [how we score](https://www.versusref.com/methodology/).

## What matters for voice agents

Weight ×5 = decisive, ×1 = relevant.

| Fact | Weight | Cartesia Ink | OpenAI gpt-4o-transcribe | Qwen3-ASR | AssemblyAI | Deepgram |
| --- | --- | --- | --- | --- | --- | --- |
| Streaming latency (vendor-claimed) | ×5 | ~88 ms (Jul 20) | n/a | n/a | ~150 ms (Jul 20) | ~300 ms (Jul 20) |
| Websocket streaming API | ×5 | ✓  Yes (Jul 20) | ✓  Yes (Jul 20) | ✓  Yes (Jul 20) | ✓  Yes (Jul 20) | ✓  Yes (Jul 20) |
| Streaming price per audio minute | ×4 | n/a | n/a | 0.005 $/audio-min (Jul 20) | 0.008 $/audio-min (Jul 20) | 0.005 $/audio-min (Jul 20) |
| Concurrency on base plan | ×3 | Free: 8, Pro: 12 concurrent STT requests (Jul 20) | Tier-based rate limits; Tier 1: 500 RPM / 10,000 TPM (Jul 20) | n/a | Async: 200+ concurrent jobs; streaming: 100 new streams/min (Jul 20) | PAYG STT: 50 REST, 150 websocket concurrent (Jul 20) |
| Custom vocabulary / keyterm boosting | ×2 | ✓  Yes (Jul 20) | ✓  Yes (Jul 20) | n/a | ✓  Yes (Jul 20) | ✓  Yes (Jul 20) |

- ×5 **Streaming latency (vendor-claimed):** The transcript delay sits inside every conversational turn; it is the floor on how fast the agent can respond.
- ×5 **Websocket streaming API:** A persistent streaming connection is the integration model voice agents are built on; batch-only APIs are a hard stop.
- ×4 **Streaming price per audio minute:** Streaming rates run higher than batch and accrue on every call minute.
- ×3 **Concurrency on base plan:** Concurrent-stream caps decide how many simultaneous calls you can serve before a bigger contract.
- ×2 **Custom vocabulary / keyterm boosting:** Boosting product names and jargon cuts the misrecognitions that derail an automated call.

## The ranking, tool by tool

| Rank | Tool | Verdict | Score | Price |
| --- | --- | --- | --- | --- |
| 1 | [Cartesia Ink](https://www.versusref.com/stt/tools/cartesia-ink/) | For voice agents, streaming latency is the top priority. | 2 of 2 points · 1 matchup | From $5/mo |
| 2 | [OpenAI gpt-4o-transcribe](https://www.versusref.com/stt/tools/gpt-4o-transcribe/) | For live voice agent use cases, real-time streaming capability is the decisive factor. | 2 of 2 points · 1 matchup | See pricing |
| 3 | [Qwen3-ASR](https://www.versusref.com/stt/tools/qwen3-asr/) (OSS) | For voice agents, a native WebSocket streaming API is essential for real-time transcription. | 2 of 2 points · 1 matchup | See pricing |
| 4 | [AssemblyAI](https://www.versusref.com/stt/tools/assemblyai/) | For voice agents, streaming capability is the decisive factor. | 9 of 12 points · 6 matchups | See pricing |
| 5 | [Deepgram](https://www.versusref.com/stt/tools/deepgram/) | For voice agents, streaming latency is the deciding factor: Deepgram publishes a 300 ms latency figure, while IBM watsonx Speech to Text publishes none. | 20 of 28 points · 14 matchups | See pricing |
| 6 | [ElevenLabs Scribe](https://www.versusref.com/stt/tools/elevenlabs-scribe/) | On the two heaviest attributes, both tools tie: streaming latency is vendor-claimed at 150 ms each, and both offer WebSocket streaming APIs. | 4 of 6 points · 3 matchups | From $6/mo |
| 7 | [Mistral Voxtral Transcribe](https://www.versusref.com/stt/tools/voxtral/) | For voice agents, streaming latency and WebSocket availability carry the most weight. | 3 of 6 points · 3 matchups | See pricing |
| 8 | [Amazon Transcribe](https://www.versusref.com/stt/tools/amazon-transcribe/) | For voice agents, a WebSocket streaming API is essential for low-latency bidirectional communication. | 2 of 4 points · 2 matchups | See pricing |
| 9 | [Azure AI Speech (STT)](https://www.versusref.com/stt/tools/azure-speech/) | For live voice agents, a native WebSocket streaming API is critical because it enables true bidirectional, low-latency communication. | 2 of 4 points · 2 matchups | From $1,600/mo |
| 10 | [Gladia](https://www.versusref.com/stt/tools/gladia/) | For voice agents, streaming latency and a WebSocket API are the two most critical requirements. | 2 of 6 points · 3 matchups | See pricing |
| 11 | [Soniox](https://www.versusref.com/stt/tools/soniox/) | On the heaviest attribute, streaming latency, Soniox claims 249 ms versus Deepgram's 300 ms, a meaningful gap when a voice agent must decide whether to interrupt a caller. | 1 of 4 points · 2 matchups | See pricing |
| 12 | [Groq (hosted Whisper)](https://www.versusref.com/stt/tools/groq-whisper/) | Ultra-fast, low-cost hosted Whisper transcription API (no realtime streaming). | 0.5 of 2 points · 1 matchup | See pricing |
| 13 | [Moonshine](https://www.versusref.com/stt/tools/moonshine/) (OSS) | On-device streaming STT for live voice interfaces, from tiny edge models to Whisper Large v3-beating accuracy. | 0.5 of 2 points · 1 matchup | See pricing |
| 14 | [NVIDIA Parakeet / Riva](https://www.versusref.com/stt/tools/nvidia-parakeet/) | Open-weights, GPU-accelerated self-hosted STT stack. | 0.5 of 4 points · 2 matchups | See pricing |
| 15 | [OpenAI Whisper (API)](https://www.versusref.com/stt/tools/whisper/) | Low-cost pay-as-you-go file transcription from OpenAI; no first-party streaming or diarization. | 1.5 of 20 points · 10 matchups | See pricing |
| 16 | [Cohere Transcribe](https://www.versusref.com/stt/tools/cohere-transcribe/) | Enterprise-grade open ASR for accurate batch transcription. | 0 of 2 points · 1 matchup | See pricing |
| 17 | [Google Cloud Speech-to-Text](https://www.versusref.com/stt/tools/google-stt/) | Hyperscaler STT API with Chirp foundation models and enterprise compliance. | 0 of 6 points · 3 matchups | See pricing |
| 18 | [xAI Grok Speech-to-Text](https://www.versusref.com/stt/tools/grok-stt/) | Low-cost hosted STT API on the Grok stack. | 0 of 2 points · 1 matchup | See pricing |
| 19 | [IBM watsonx Speech to Text](https://www.versusref.com/stt/tools/ibm-watson-stt/) | Enterprise cloud STT API. | 0 of 2 points · 1 matchup | See pricing |
| 20 | [Rev AI](https://www.versusref.com/stt/tools/rev/) | Transcription-heritage STT API with low per-hour pricing and open (non-commercial) Reverb models. | 0 of 4 points · 2 matchups | See pricing |
| 21 | [Speechmatics](https://www.versusref.com/stt/tools/speechmatics/) | Accuracy-first enterprise STT with flexible deployment (SaaS, container, on-prem). | 0 of 4 points · 2 matchups | See pricing |

### 1. Cartesia Ink

For voice agents, streaming latency is the top priority. [Full Cartesia Ink vs Deepgram verdict](https://www.versusref.com/stt/cartesia-ink-vs-deepgram/)

### 2. OpenAI gpt-4o-transcribe

For live voice agent use cases, real-time streaming capability is the decisive factor. [Full OpenAI gpt-4o-transcribe vs OpenAI Whisper (API) verdict](https://www.versusref.com/stt/gpt-4o-transcribe-vs-whisper/)

### 3. Qwen3-ASR

For voice agents, a native WebSocket streaming API is essential for real-time transcription. [Full Qwen3-ASR vs OpenAI Whisper (API) verdict](https://www.versusref.com/stt/qwen3-asr-vs-whisper/)

### 4. AssemblyAI

For voice agents, streaming capability is the decisive factor. [Full AssemblyAI vs OpenAI Whisper (API) verdict](https://www.versusref.com/stt/assemblyai-vs-whisper/)

Both tools offer WebSocket streaming APIs, so that factor cancels out. [Full AssemblyAI vs Speechmatics verdict](https://www.versusref.com/stt/assemblyai-vs-speechmatics/)

For voice agents, streaming latency is the top-weighted factor. [Full AssemblyAI vs Soniox verdict](https://www.versusref.com/stt/assemblyai-vs-soniox/)

For voice agents, streaming latency is the top priority. [Full AssemblyAI vs Rev AI verdict](https://www.versusref.com/stt/assemblyai-vs-rev/)

For voice agents, streaming latency is the decisive attribute. [Full AssemblyAI vs Gladia verdict](https://www.versusref.com/stt/assemblyai-vs-gladia/)

### 5. Deepgram

For voice agents, streaming latency is the deciding factor: Deepgram publishes a 300 ms latency figure, while IBM watsonx Speech to Text publishes none. [Full Deepgram vs IBM watsonx Speech to Text verdict](https://www.versusref.com/stt/deepgram-vs-ibm-watson-stt/)

For voice agents, streaming latency is the top-weighted factor. [Full Deepgram vs Amazon Transcribe verdict](https://www.versusref.com/stt/amazon-transcribe-vs-deepgram/)

For voice agents, streaming latency is the top-weighted factor. [Full Deepgram vs Azure AI Speech (STT) verdict](https://www.versusref.com/stt/azure-speech-vs-deepgram/)

For voice agents, streaming latency is the top factor. [Full Deepgram vs Rev AI verdict](https://www.versusref.com/stt/deepgram-vs-rev/)

For voice agents, streaming latency and websocket support are the two heaviest factors, and Deepgram dominates both. [Full Deepgram vs NVIDIA Parakeet / Riva verdict](https://www.versusref.com/stt/deepgram-vs-nvidia-parakeet/)

Both tools offer websocket streaming and custom vocabulary, so those attributes are tied. [Full Deepgram vs xAI Grok Speech-to-Text verdict](https://www.versusref.com/stt/deepgram-vs-grok-stt/)

For voice agents, streaming capability and latency are the top two factors, and Deepgram leads decisively on both. [Full Deepgram vs OpenAI Whisper (API) verdict](https://www.versusref.com/stt/deepgram-vs-whisper/)

For voice agents, streaming latency is the decisive factor. [Full Deepgram vs Speechmatics verdict](https://www.versusref.com/stt/deepgram-vs-speechmatics/)

For voice agents, streaming latency and websocket support are the heaviest factors. [Full Deepgram vs Google Cloud Speech-to-Text verdict](https://www.versusref.com/stt/deepgram-vs-google-stt/)

Both tools match on streaming latency at 300 ms and both offer websocket streaming APIs, so the top two attributes are tied. [Full Deepgram vs Gladia verdict](https://www.versusref.com/stt/deepgram-vs-gladia/)

For voice agents, streaming latency and websocket support are the two most critical attributes, both weighted 5 out of 5. [Full Deepgram vs Cohere Transcribe verdict](https://www.versusref.com/stt/cohere-transcribe-vs-deepgram/)

### 6. ElevenLabs Scribe

On the two heaviest attributes, both tools tie: streaming latency is vendor-claimed at 150 ms each, and both offer WebSocket streaming APIs. [Full ElevenLabs Scribe vs AssemblyAI verdict](https://www.versusref.com/stt/assemblyai-vs-elevenlabs-scribe/)

For voice agents, streaming latency and WebSocket support are the two most critical attributes. [Full ElevenLabs Scribe vs OpenAI Whisper (API) verdict](https://www.versusref.com/stt/elevenlabs-scribe-vs-whisper/)

On the two heaviest attributes, ElevenLabs Scribe leads on streaming latency at 150 ms versus Mistral Voxtral Transcribe at 200 ms, and both tools offer a WebSocket streaming API. [Full ElevenLabs Scribe vs Mistral Voxtral Transcribe verdict](https://www.versusref.com/stt/elevenlabs-scribe-vs-voxtral/)

### 7. Mistral Voxtral Transcribe

For voice agents, streaming latency and WebSocket availability carry the most weight. [Full Mistral Voxtral Transcribe vs Deepgram verdict](https://www.versusref.com/stt/deepgram-vs-voxtral/)

For live voice agents, streaming capability and latency are decisive. [Full Mistral Voxtral Transcribe vs OpenAI Whisper (API) verdict](https://www.versusref.com/stt/voxtral-vs-whisper/)

### 8. Amazon Transcribe

For voice agents, a WebSocket streaming API is essential for low-latency bidirectional communication. [Full Amazon Transcribe vs Google Cloud Speech-to-Text verdict](https://www.versusref.com/stt/amazon-transcribe-vs-google-stt/)

### 9. Azure AI Speech (STT)

For live voice agents, a native WebSocket streaming API is critical because it enables true bidirectional, low-latency communication. [Full Azure AI Speech (STT) vs Google Cloud Speech-to-Text verdict](https://www.versusref.com/stt/azure-speech-vs-google-stt/)

### 10. Gladia

For voice agents, streaming latency and a WebSocket API are the two most critical requirements. [Full Gladia vs OpenAI Whisper (API) verdict](https://www.versusref.com/stt/gladia-vs-whisper/)

### 11. Soniox

On the heaviest attribute, streaming latency, Soniox claims 249 ms versus Deepgram's 300 ms, a meaningful gap when a voice agent must decide whether to interrupt a caller. [Full Soniox vs Deepgram verdict](https://www.versusref.com/stt/deepgram-vs-soniox/)

### 12. Groq (hosted Whisper)

Ultra-fast, low-cost hosted Whisper transcription API (no realtime streaming). No won verdicts for this use case yet; it ranks on ties and near-misses.

### 13. Moonshine

On-device streaming STT for live voice interfaces, from tiny edge models to Whisper Large v3-beating accuracy. No won verdicts for this use case yet; it ranks on ties and near-misses.

### 14. NVIDIA Parakeet / Riva

Open-weights, GPU-accelerated self-hosted STT stack. No won verdicts for this use case yet; it ranks on ties and near-misses.

### 15. OpenAI Whisper (API)

Low-cost pay-as-you-go file transcription from OpenAI; no first-party streaming or diarization. No won verdicts for this use case yet; it ranks on ties and near-misses.

### 16. Cohere Transcribe

Enterprise-grade open ASR for accurate batch transcription. No won verdicts for this use case yet; it ranks on ties and near-misses.

### 17. Google Cloud Speech-to-Text

Hyperscaler STT API with Chirp foundation models and enterprise compliance. No won verdicts for this use case yet; it ranks on ties and near-misses.

### 18. xAI Grok Speech-to-Text

Low-cost hosted STT API on the Grok stack. No won verdicts for this use case yet; it ranks on ties and near-misses.

### 19. IBM watsonx Speech to Text

Enterprise cloud STT API. No won verdicts for this use case yet; it ranks on ties and near-misses.

### 20. Rev AI

Transcription-heritage STT API with low per-hour pricing and open (non-commercial) Reverb models. No won verdicts for this use case yet; it ranks on ties and near-misses.

### 21. Speechmatics

Accuracy-first enterprise STT with flexible deployment (SaaS, container, on-prem). No won verdicts for this use case yet; it ranks on ties and near-misses.

Source: https://www.versusref.com/stt/best/voice-agents/
