Best Speech-to-text APIs for Self-Hosted (2026)
For self-hosted, Mistral Voxtral Transcribe is our pick: For self-hosted deployment, the two heaviest attributes are the self-host option and an open-weights license. Running an open-weight speech model on your own hardware: license reality, hardware needs, and whether the project is still alive. Below is the full ranking and the tradeoffs, or read how we score.
If you sign up through links on this page, vsref may earn a commission; programs exist on both sides of most comparisons, and commissions never change verdicts. How we make money
Reviewed by vsref Editorialfacts verified Jul 24, 2026Methodology →
What matters for self-hosted
Weight ×5 = decisive, ×1 = relevant| Fact | Mistral Voxtral Transcribe | NVIDIA Parakeet / Riva | Cohere Transcribe | Groq (hosted Whisper) | Moonshine |
|---|---|---|---|---|---|
| Self-host / on-prem option×5 | ✓ YesJul 20 | ✓ YesJul 20 | ✓ YesJul 20 | ✓ YesJul 20 | ✓ YesJul 20 |
| Model weights license×5 | Apache-2.0Jul 20 | CC-BY-4.0Jul 20 | Apache 2.0Jul 20 | n/a | MIT (English models); others non-commercialJul 20 |
| Hardware to self-host×4 | n/a | NVIDIA GPU (T4 to H100); min 2GB RAMJul 20 | Small GPU; exact VRAM not publishedJul 20 | n/a | CPU-only; desktop, mobile, Raspberry Pi, MCUs, DSPsJul 20 |
| Project maintenance status×3 | n/a | Active - NeMo v2.7.3 released 2026-04-23Jul 20 | Active - released 2026-03; Arabic 2026-07Jul 20 | n/a | Active - v0.0.69 released 2026-07-16Jul 20 |
| Model size (parameters)×2 | n/a | Parakeet TDT 0.6B (600M); Canary 1B v2 (978M)Jul 20 | 2BJul 20 | n/a | 26M-245M (Tiny to Medium Streaming)Jul 20 |
The ranking, tool by tool
For self-hosted deployment, the two heaviest attributes are the self-host option and an open-weights license. Full Mistral Voxtral Transcribe vs Deepgram verdict →
For self-hosted deployment, the two highest-weighted attributes are the self-host option and the model weights license. Full Mistral Voxtral Transcribe vs OpenAI Whisper (API) verdict →
For self-hosted deployment, the two most critical factors are whether self-hosting is possible and the model weights license. Full Mistral Voxtral Transcribe vs ElevenLabs Scribe verdict →
For self-hosted deployment, the two heaviest attributes are the self-host option and the model weights license. Full NVIDIA Parakeet / Riva vs Deepgram verdict →
For self-hosted deployment, NVIDIA Parakeet / Riva wins on every attribute that matters. Full NVIDIA Parakeet / Riva vs OpenAI Whisper (API) verdict →
For self-hosted deployment, the model weights license is decisive. Full Cohere Transcribe vs Deepgram verdict →
For self-hosted deployment, the two heaviest attributes are the self-host option and the model weights license. Full Groq (hosted Whisper) vs OpenAI Whisper (API) verdict →
Moonshine is purpose-built for self-hosted deployment, while OpenAI Whisper (API) explicitly offers no self-host option. Full Moonshine vs OpenAI Whisper (API) verdict →
For self-hosted deployment, Qwen3-ASR leads on every attribute that matters. Full Qwen3-ASR vs OpenAI Whisper (API) verdict →
For a self-hosted deployment, the single most critical factor is whether a vendor actually offers a self-host or on-prem option. Full AssemblyAI vs Soniox verdict →
AssemblyAI explicitly supports self-hosted and on-premises deployment, which is the single most critical requirement for this use case. Full AssemblyAI vs Rev AI verdict →
For self-hosted deployment, the single most critical factor is whether a vendor actually offers a self-host or on-prem option. Full AssemblyAI vs ElevenLabs Scribe verdict →
For self-hosted deployment on your own hardware, both heavily weighted attributes favor OpenAI Whisper. Full OpenAI Whisper (API) vs AssemblyAI verdict →
For self-hosted deployment on your own hardware, the two most important factors are the self-host option and the model weights license. Full OpenAI Whisper (API) vs Gladia verdict →
For self-hosted deployment, the two heaviest attributes are whether a tool can be self-hosted and whether it ships with an open-weights license. Full OpenAI Whisper (API) vs ElevenLabs Scribe verdict →
For self-hosted deployment, the two heaviest attributes are the self-host option and model weights license, each weighted 5 out of 5. Full OpenAI Whisper (API) vs Deepgram verdict →
For self-hosted deployment, Deepgram is the clear winner. Full Deepgram vs ElevenLabs Scribe verdict →
For a self-hosted deployment, the single most important question is whether the vendor allows on-premises installation at all. Full Deepgram vs Amazon Transcribe verdict →
For a self-hosted deployment, the single most critical requirement is whether the vendor supports running the model on your own hardware. Full Deepgram vs xAI Grok Speech-to-Text verdict →
For self-hosted deployments, the single most critical requirement is whether a vendor offers an on-premises option. Full Deepgram vs Soniox verdict →
Enterprise-grade STT inside the Azure cloud ecosystem. No won verdicts for this use case yet; it ranks on ties and near-misses.
Hyperscaler STT API with Chirp foundation models and enterprise compliance. No won verdicts for this use case yet; it ranks on ties and near-misses.
Accuracy-first enterprise STT with flexible deployment (SaaS, container, on-prem). No won verdicts for this use case yet; it ranks on ties and near-misses.
Streaming STT for voice agents with native turn detection. No won verdicts for this use case yet; it ranks on ties and near-misses.
Flagship GPT-4o based transcription API from OpenAI. No won verdicts for this use case yet; it ranks on ties and near-misses.
Ultra-low-latency multilingual STT for voice agents. No won verdicts for this use case yet; it ranks on ties and near-misses.
EU-based real-time and batch STT API built on the Solaria models. No won verdicts for this use case yet; it ranks on ties and near-misses.
AWS-native STT API with deep AWS ecosystem integration and compliance coverage. No won verdicts for this use case yet; it ranks on ties and near-misses.
Accuracy-led STT API from the leading AI audio company. No won verdicts for this use case yet; it ranks on ties and near-misses.
Low-cost hosted STT API on the Grok stack. No won verdicts for this use case yet; it ranks on ties and near-misses.
Transcription-heritage STT API with low per-hour pricing and open (non-commercial) Reverb models. No won verdicts for this use case yet; it ranks on ties and near-misses.
Ultra-low-cost multilingual STT + real-time translation API. No won verdicts for this use case yet; it ranks on ties and near-misses.