vsref

Best Speech-to-text APIs for Self-Hosted (2026)

For self-hosted, Mistral Voxtral Transcribe is our pick: For self-hosted deployment, the two heaviest attributes are the self-host option and an open-weights license. Running an open-weight speech model on your own hardware: license reality, hardware needs, and whether the project is still alive. Below is the full ranking and the tradeoffs, or read how we score.

If you sign up through links on this page, vsref may earn a commission; programs exist on both sides of most comparisons, and commissions never change verdicts. How we make money

Reviewed by vsref Editorialfacts verified Jul 24, 2026Methodology →

Low-cost EU-based transcription API with Apache-2.0 open-weight models6 of 6 points · 3 matchups
Open-weights, GPU-accelerated self-hosted STT stack4 of 4 points · 2 matchups
Enterprise-grade open ASR for accurate batch transcription2 of 2 points · 1 matchup

What matters for self-hosted

Weighted attribute comparison for Self-Hosted
FactMistral Voxtral TranscribeNVIDIA Parakeet / RivaCohere TranscribeGroq (hosted Whisper)Moonshine
Self-host / on-prem option×5✓ YesJul 20✓ YesJul 20✓ YesJul 20✓ YesJul 20✓ YesJul 20
Model weights license×5Apache-2.0Jul 20CC-BY-4.0Jul 20Apache 2.0Jul 20n/aMIT (English models); others non-commercialJul 20
Hardware to self-host×4n/aNVIDIA GPU (T4 to H100); min 2GB RAMJul 20Small GPU; exact VRAM not publishedJul 20n/aCPU-only; desktop, mobile, Raspberry Pi, MCUs, DSPsJul 20
Project maintenance status×3n/aActive - NeMo v2.7.3 released 2026-04-23Jul 20Active - released 2026-03; Arabic 2026-07Jul 20n/aActive - v0.0.69 released 2026-07-16Jul 20
Model size (parameters)×2n/aParakeet TDT 0.6B (600M); Canary 1B v2 (978M)Jul 202BJul 20n/a26M-245M (Tiny to Medium Streaming)Jul 20
Swipe → to see every tool column.
×5 Self-host / on-prem option: This page only ranks what you can actually run yourself.×5 Model weights license: The license decides commercial viability: MIT and Apache ship products, non-commercial licenses do not.×4 Hardware to self-host: CPU-capable small models and GPU-hungry large ones are different budgets entirely.×3 Project maintenance status: A dormant repo means you own every future bug and compatibility break.×2 Model size (parameters): Parameter count is the quickest proxy for the speed/accuracy trade-off on your hardware.

The ranking, tool by tool

For self-hosted deployment, the two heaviest attributes are the self-host option and an open-weights license.

For self-hosted deployment, the two heaviest attributes are the self-host option and an open-weights license. Full Mistral Voxtral Transcribe vs Deepgram verdict →

For self-hosted deployment, the two highest-weighted attributes are the self-host option and the model weights license. Full Mistral Voxtral Transcribe vs OpenAI Whisper (API) verdict →

For self-hosted deployment, the two most critical factors are whether self-hosting is possible and the model weights license. Full Mistral Voxtral Transcribe vs ElevenLabs Scribe verdict →

For self-hosted deployment, the two heaviest attributes are the self-host option and the model weights license.

For self-hosted deployment, the two heaviest attributes are the self-host option and the model weights license. Full NVIDIA Parakeet / Riva vs Deepgram verdict →

For self-hosted deployment, NVIDIA Parakeet / Riva wins on every attribute that matters. Full NVIDIA Parakeet / Riva vs OpenAI Whisper (API) verdict →

For self-hosted deployment, the model weights license is decisive.

For self-hosted deployment, the model weights license is decisive. Full Cohere Transcribe vs Deepgram verdict →

For self-hosted deployment, the two heaviest attributes are the self-host option and the model weights license.

For self-hosted deployment, the two heaviest attributes are the self-host option and the model weights license. Full Groq (hosted Whisper) vs OpenAI Whisper (API) verdict →

Moonshine is purpose-built for self-hosted deployment, while OpenAI Whisper (API) explicitly offers no self-host option.
See pricingWebsite →

Moonshine is purpose-built for self-hosted deployment, while OpenAI Whisper (API) explicitly offers no self-host option. Full Moonshine vs OpenAI Whisper (API) verdict →

For self-hosted deployment, Qwen3-ASR leads on every attribute that matters.
See pricingWebsite →

For self-hosted deployment, Qwen3-ASR leads on every attribute that matters. Full Qwen3-ASR vs OpenAI Whisper (API) verdict →

For a self-hosted deployment, the single most critical factor is whether a vendor actually offers a self-host or on-prem option.

For a self-hosted deployment, the single most critical factor is whether a vendor actually offers a self-host or on-prem option. Full AssemblyAI vs Soniox verdict →

AssemblyAI explicitly supports self-hosted and on-premises deployment, which is the single most critical requirement for this use case. Full AssemblyAI vs Rev AI verdict →

For self-hosted deployment, the single most critical factor is whether a vendor actually offers a self-host or on-prem option. Full AssemblyAI vs ElevenLabs Scribe verdict →

For self-hosted deployment on your own hardware, both heavily weighted attributes favor OpenAI Whisper.

For self-hosted deployment on your own hardware, both heavily weighted attributes favor OpenAI Whisper. Full OpenAI Whisper (API) vs AssemblyAI verdict →

For self-hosted deployment on your own hardware, the two most important factors are the self-host option and the model weights license. Full OpenAI Whisper (API) vs Gladia verdict →

For self-hosted deployment, the two heaviest attributes are whether a tool can be self-hosted and whether it ships with an open-weights license. Full OpenAI Whisper (API) vs ElevenLabs Scribe verdict →

For self-hosted deployment, the two heaviest attributes are the self-host option and model weights license, each weighted 5 out of 5. Full OpenAI Whisper (API) vs Deepgram verdict →

For self-hosted deployment, Deepgram is the clear winner.
See pricingTry Deepgram

For self-hosted deployment, Deepgram is the clear winner. Full Deepgram vs ElevenLabs Scribe verdict →

For a self-hosted deployment, the single most important question is whether the vendor allows on-premises installation at all. Full Deepgram vs Amazon Transcribe verdict →

For a self-hosted deployment, the single most critical requirement is whether the vendor supports running the model on your own hardware. Full Deepgram vs xAI Grok Speech-to-Text verdict →

For self-hosted deployments, the single most critical requirement is whether a vendor offers an on-premises option. Full Deepgram vs Soniox verdict →

Enterprise-grade STT inside the Azure cloud ecosystem.

Enterprise-grade STT inside the Azure cloud ecosystem. No won verdicts for this use case yet; it ranks on ties and near-misses.

Hyperscaler STT API with Chirp foundation models and enterprise compliance.

Hyperscaler STT API with Chirp foundation models and enterprise compliance. No won verdicts for this use case yet; it ranks on ties and near-misses.

Accuracy-first enterprise STT with flexible deployment (SaaS, container, on-prem).

Accuracy-first enterprise STT with flexible deployment (SaaS, container, on-prem). No won verdicts for this use case yet; it ranks on ties and near-misses.

Streaming STT for voice agents with native turn detection.

Streaming STT for voice agents with native turn detection. No won verdicts for this use case yet; it ranks on ties and near-misses.

Flagship GPT-4o based transcription API from OpenAI.

Flagship GPT-4o based transcription API from OpenAI. No won verdicts for this use case yet; it ranks on ties and near-misses.

Ultra-low-latency multilingual STT for voice agents.

Ultra-low-latency multilingual STT for voice agents. No won verdicts for this use case yet; it ranks on ties and near-misses.

EU-based real-time and batch STT API built on the Solaria models.
See pricingTry Gladia

EU-based real-time and batch STT API built on the Solaria models. No won verdicts for this use case yet; it ranks on ties and near-misses.

AWS-native STT API with deep AWS ecosystem integration and compliance coverage.

AWS-native STT API with deep AWS ecosystem integration and compliance coverage. No won verdicts for this use case yet; it ranks on ties and near-misses.

Accuracy-led STT API from the leading AI audio company.

Accuracy-led STT API from the leading AI audio company. No won verdicts for this use case yet; it ranks on ties and near-misses.

Low-cost hosted STT API on the Grok stack.

Low-cost hosted STT API on the Grok stack. No won verdicts for this use case yet; it ranks on ties and near-misses.

Transcription-heritage STT API with low per-hour pricing and open (non-commercial) Reverb models.
See pricingTry Rev AI

Transcription-heritage STT API with low per-hour pricing and open (non-commercial) Reverb models. No won verdicts for this use case yet; it ranks on ties and near-misses.

Ultra-low-cost multilingual STT + real-time translation API.
See pricingTry Soniox

Ultra-low-cost multilingual STT + real-time translation API. No won verdicts for this use case yet; it ranks on ties and near-misses.

More Speech-to-text APIs buyer guides