# Best Speech-to-text APIs for Self-Hosted (2026)

> The best Speech-to-text APIs platforms for self-hosted: Mistral Voxtral Transcribe leads, for running an open-weight speech model on your own hardware: license.

For self-hosted, **Mistral Voxtral Transcribe** is our pick: For self-hosted deployment, the two heaviest attributes are the self-host option and an open-weights license. Running an open-weight speech model on your own hardware: license reality, hardware needs, and whether the project is still alive. Below is the full ranking and the tradeoffs, or read [how we score](https://www.versusref.com/methodology/).

## What matters for self-hosted

Weight ×5 = decisive, ×1 = relevant.

| Fact | Weight | Mistral Voxtral Transcribe | NVIDIA Parakeet / Riva | Cohere Transcribe | Groq (hosted Whisper) | Moonshine |
| --- | --- | --- | --- | --- | --- | --- |
| Self-host / on-prem option | ×5 | ✓  Yes (Jul 20) | ✓  Yes (Jul 20) | ✓  Yes (Jul 20) | ✓  Yes (Jul 20) | ✓  Yes (Jul 20) |
| Model weights license | ×5 | Apache-2.0 (Jul 20) | CC-BY-4.0 (Jul 20) | Apache 2.0 (Jul 20) | n/a | MIT (English models); others non-commercial (Jul 20) |
| Hardware to self-host | ×4 | n/a | NVIDIA GPU (T4 to H100); min 2GB RAM (Jul 20) | Small GPU; exact VRAM not published (Jul 20) | n/a | CPU-only; desktop, mobile, Raspberry Pi, MCUs, DSPs (Jul 20) |
| Project maintenance status | ×3 | n/a | Active - NeMo v2.7.3 released 2026-04-23 (Jul 20) | Active - released 2026-03; Arabic 2026-07 (Jul 20) | n/a | Active - v0.0.69 released 2026-07-16 (Jul 20) |
| Model size (parameters) | ×2 | n/a | Parakeet TDT 0.6B (600M); Canary 1B v2 (978M) (Jul 20) | 2B (Jul 20) | n/a | 26M-245M (Tiny to Medium Streaming) (Jul 20) |

- ×5 **Self-host / on-prem option:** This page only ranks what you can actually run yourself.
- ×5 **Model weights license:** The license decides commercial viability: MIT and Apache ship products, non-commercial licenses do not.
- ×4 **Hardware to self-host:** CPU-capable small models and GPU-hungry large ones are different budgets entirely.
- ×3 **Project maintenance status:** A dormant repo means you own every future bug and compatibility break.
- ×2 **Model size (parameters):** Parameter count is the quickest proxy for the speed/accuracy trade-off on your hardware.

## The ranking, tool by tool

| Rank | Tool | Verdict | Score | Price |
| --- | --- | --- | --- | --- |
| 1 | [Mistral Voxtral Transcribe](https://www.versusref.com/stt/tools/voxtral/) | For self-hosted deployment, the two heaviest attributes are the self-host option and an open-weights license. | 6 of 6 points · 3 matchups | See pricing |
| 2 | [NVIDIA Parakeet / Riva](https://www.versusref.com/stt/tools/nvidia-parakeet/) | For self-hosted deployment, the two heaviest attributes are the self-host option and the model weights license. | 4 of 4 points · 2 matchups | See pricing |
| 3 | [Cohere Transcribe](https://www.versusref.com/stt/tools/cohere-transcribe/) | For self-hosted deployment, the model weights license is decisive. | 2 of 2 points · 1 matchup | See pricing |
| 4 | [Groq (hosted Whisper)](https://www.versusref.com/stt/tools/groq-whisper/) | For self-hosted deployment, the two heaviest attributes are the self-host option and the model weights license. | 2 of 2 points · 1 matchup | See pricing |
| 5 | [Moonshine](https://www.versusref.com/stt/tools/moonshine/) (OSS) | Moonshine is purpose-built for self-hosted deployment, while OpenAI Whisper (API) explicitly offers no self-host option. | 2 of 2 points · 1 matchup | See pricing |
| 6 | [Qwen3-ASR](https://www.versusref.com/stt/tools/qwen3-asr/) (OSS) | For self-hosted deployment, Qwen3-ASR leads on every attribute that matters. | 2 of 2 points · 1 matchup | See pricing |
| 7 | [AssemblyAI](https://www.versusref.com/stt/tools/assemblyai/) | For a self-hosted deployment, the single most critical factor is whether a vendor actually offers a self-host or on-prem option. | 7.5 of 14 points · 7 matchups | See pricing |
| 8 | [OpenAI Whisper (API)](https://www.versusref.com/stt/tools/whisper/) | For self-hosted deployment on your own hardware, both heavily weighted attributes favor OpenAI Whisper. | 8.5 of 20 points · 10 matchups | See pricing |
| 9 | [Deepgram](https://www.versusref.com/stt/tools/deepgram/) | For self-hosted deployment, Deepgram is the clear winner. | 12 of 32 points · 16 matchups | See pricing |
| 10 | [Azure AI Speech (STT)](https://www.versusref.com/stt/tools/azure-speech/) | Enterprise-grade STT inside the Azure cloud ecosystem. | 1 of 4 points · 2 matchups | From $1,600/mo |
| 11 | [Google Cloud Speech-to-Text](https://www.versusref.com/stt/tools/google-stt/) | Hyperscaler STT API with Chirp foundation models and enterprise compliance. | 1 of 4 points · 2 matchups | See pricing |
| 12 | [Speechmatics](https://www.versusref.com/stt/tools/speechmatics/) | Accuracy-first enterprise STT with flexible deployment (SaaS, container, on-prem). | 1 of 4 points · 2 matchups | See pricing |
| 13 | [Cartesia Ink](https://www.versusref.com/stt/tools/cartesia-ink/) | Streaming STT for voice agents with native turn detection. | 0.5 of 2 points · 1 matchup | From $5/mo |
| 14 | [OpenAI gpt-4o-transcribe](https://www.versusref.com/stt/tools/gpt-4o-transcribe/) | Flagship GPT-4o based transcription API from OpenAI. | 0.5 of 2 points · 1 matchup | See pricing |
| 15 | [IBM watsonx Speech to Text](https://www.versusref.com/stt/tools/ibm-watson-stt/) | Enterprise cloud STT API. | 0.5 of 2 points · 1 matchup | See pricing |
| 16 | [Smallest.ai Pulse](https://www.versusref.com/stt/tools/smallest-pulse/) | Ultra-low-latency multilingual STT for voice agents. | 0.5 of 2 points · 1 matchup | See pricing |
| 17 | [Gladia](https://www.versusref.com/stt/tools/gladia/) | EU-based real-time and batch STT API built on the Solaria models. | 1 of 6 points · 3 matchups | See pricing |
| 18 | [Amazon Transcribe](https://www.versusref.com/stt/tools/amazon-transcribe/) | AWS-native STT API with deep AWS ecosystem integration and compliance coverage. | 0 of 2 points · 1 matchup | See pricing |
| 19 | [ElevenLabs Scribe](https://www.versusref.com/stt/tools/elevenlabs-scribe/) | Accuracy-led STT API from the leading AI audio company. | 0 of 8 points · 4 matchups | From $6/mo |
| 20 | [xAI Grok Speech-to-Text](https://www.versusref.com/stt/tools/grok-stt/) | Low-cost hosted STT API on the Grok stack. | 0 of 2 points · 1 matchup | See pricing |
| 21 | [Rev AI](https://www.versusref.com/stt/tools/rev/) | Transcription-heritage STT API with low per-hour pricing and open (non-commercial) Reverb models. | 0 of 2 points · 1 matchup | See pricing |
| 22 | [Soniox](https://www.versusref.com/stt/tools/soniox/) | Ultra-low-cost multilingual STT + real-time translation API. | 0 of 4 points · 2 matchups | See pricing |

### 1. Mistral Voxtral Transcribe

For self-hosted deployment, the two heaviest attributes are the self-host option and an open-weights license. [Full Mistral Voxtral Transcribe vs Deepgram verdict](https://www.versusref.com/stt/deepgram-vs-voxtral/)

For self-hosted deployment, the two highest-weighted attributes are the self-host option and the model weights license. [Full Mistral Voxtral Transcribe vs OpenAI Whisper (API) verdict](https://www.versusref.com/stt/voxtral-vs-whisper/)

For self-hosted deployment, the two most critical factors are whether self-hosting is possible and the model weights license. [Full Mistral Voxtral Transcribe vs ElevenLabs Scribe verdict](https://www.versusref.com/stt/elevenlabs-scribe-vs-voxtral/)

### 2. NVIDIA Parakeet / Riva

For self-hosted deployment, the two heaviest attributes are the self-host option and the model weights license. [Full NVIDIA Parakeet / Riva vs Deepgram verdict](https://www.versusref.com/stt/deepgram-vs-nvidia-parakeet/)

For self-hosted deployment, NVIDIA Parakeet / Riva wins on every attribute that matters. [Full NVIDIA Parakeet / Riva vs OpenAI Whisper (API) verdict](https://www.versusref.com/stt/nvidia-parakeet-vs-whisper/)

### 3. Cohere Transcribe

For self-hosted deployment, the model weights license is decisive. [Full Cohere Transcribe vs Deepgram verdict](https://www.versusref.com/stt/cohere-transcribe-vs-deepgram/)

### 4. Groq (hosted Whisper)

For self-hosted deployment, the two heaviest attributes are the self-host option and the model weights license. [Full Groq (hosted Whisper) vs OpenAI Whisper (API) verdict](https://www.versusref.com/stt/groq-whisper-vs-whisper/)

### 5. Moonshine

Moonshine is purpose-built for self-hosted deployment, while OpenAI Whisper (API) explicitly offers no self-host option. [Full Moonshine vs OpenAI Whisper (API) verdict](https://www.versusref.com/stt/moonshine-vs-whisper/)

### 6. Qwen3-ASR

For self-hosted deployment, Qwen3-ASR leads on every attribute that matters. [Full Qwen3-ASR vs OpenAI Whisper (API) verdict](https://www.versusref.com/stt/qwen3-asr-vs-whisper/)

### 7. AssemblyAI

For a self-hosted deployment, the single most critical factor is whether a vendor actually offers a self-host or on-prem option. [Full AssemblyAI vs Soniox verdict](https://www.versusref.com/stt/assemblyai-vs-soniox/)

AssemblyAI explicitly supports self-hosted and on-premises deployment, which is the single most critical requirement for this use case. [Full AssemblyAI vs Rev AI verdict](https://www.versusref.com/stt/assemblyai-vs-rev/)

For self-hosted deployment, the single most critical factor is whether a vendor actually offers a self-host or on-prem option. [Full AssemblyAI vs ElevenLabs Scribe verdict](https://www.versusref.com/stt/assemblyai-vs-elevenlabs-scribe/)

### 8. OpenAI Whisper (API)

For self-hosted deployment on your own hardware, both heavily weighted attributes favor OpenAI Whisper. [Full OpenAI Whisper (API) vs AssemblyAI verdict](https://www.versusref.com/stt/assemblyai-vs-whisper/)

For self-hosted deployment on your own hardware, the two most important factors are the self-host option and the model weights license. [Full OpenAI Whisper (API) vs Gladia verdict](https://www.versusref.com/stt/gladia-vs-whisper/)

For self-hosted deployment, the two heaviest attributes are whether a tool can be self-hosted and whether it ships with an open-weights license. [Full OpenAI Whisper (API) vs ElevenLabs Scribe verdict](https://www.versusref.com/stt/elevenlabs-scribe-vs-whisper/)

For self-hosted deployment, the two heaviest attributes are the self-host option and model weights license, each weighted 5 out of 5. [Full OpenAI Whisper (API) vs Deepgram verdict](https://www.versusref.com/stt/deepgram-vs-whisper/)

### 9. Deepgram

For self-hosted deployment, Deepgram is the clear winner. [Full Deepgram vs ElevenLabs Scribe verdict](https://www.versusref.com/stt/deepgram-vs-elevenlabs-scribe/)

For a self-hosted deployment, the single most important question is whether the vendor allows on-premises installation at all. [Full Deepgram vs Amazon Transcribe verdict](https://www.versusref.com/stt/amazon-transcribe-vs-deepgram/)

For a self-hosted deployment, the single most critical requirement is whether the vendor supports running the model on your own hardware. [Full Deepgram vs xAI Grok Speech-to-Text verdict](https://www.versusref.com/stt/deepgram-vs-grok-stt/)

For self-hosted deployments, the single most critical requirement is whether a vendor offers an on-premises option. [Full Deepgram vs Soniox verdict](https://www.versusref.com/stt/deepgram-vs-soniox/)

### 10. Azure AI Speech (STT)

Enterprise-grade STT inside the Azure cloud ecosystem. No won verdicts for this use case yet; it ranks on ties and near-misses.

### 11. Google Cloud Speech-to-Text

Hyperscaler STT API with Chirp foundation models and enterprise compliance. No won verdicts for this use case yet; it ranks on ties and near-misses.

### 12. Speechmatics

Accuracy-first enterprise STT with flexible deployment (SaaS, container, on-prem). No won verdicts for this use case yet; it ranks on ties and near-misses.

### 13. Cartesia Ink

Streaming STT for voice agents with native turn detection. No won verdicts for this use case yet; it ranks on ties and near-misses.

### 14. OpenAI gpt-4o-transcribe

Flagship GPT-4o based transcription API from OpenAI. No won verdicts for this use case yet; it ranks on ties and near-misses.

### 15. IBM watsonx Speech to Text

Enterprise cloud STT API. No won verdicts for this use case yet; it ranks on ties and near-misses.

### 16. Smallest.ai Pulse

Ultra-low-latency multilingual STT for voice agents. No won verdicts for this use case yet; it ranks on ties and near-misses.

### 17. Gladia

EU-based real-time and batch STT API built on the Solaria models. No won verdicts for this use case yet; it ranks on ties and near-misses.

### 18. Amazon Transcribe

AWS-native STT API with deep AWS ecosystem integration and compliance coverage. No won verdicts for this use case yet; it ranks on ties and near-misses.

### 19. ElevenLabs Scribe

Accuracy-led STT API from the leading AI audio company. No won verdicts for this use case yet; it ranks on ties and near-misses.

### 20. xAI Grok Speech-to-Text

Low-cost hosted STT API on the Grok stack. No won verdicts for this use case yet; it ranks on ties and near-misses.

### 21. Rev AI

Transcription-heritage STT API with low per-hour pricing and open (non-commercial) Reverb models. No won verdicts for this use case yet; it ranks on ties and near-misses.

### 22. Soniox

Ultra-low-cost multilingual STT + real-time translation API. No won verdicts for this use case yet; it ranks on ties and near-misses.

Source: https://www.versusref.com/stt/best/self-hosted/
