vsref

Best Speech-to-text APIs for Dictation (2026)

For dictation, Moonshine is our pick: For voice typing on your own computer, local processing is the dominant concern because users want audio to stay on device. Voice typing on your own computer: which apps work where, whether audio leaves the device, and what the pricing model really is. Below is the full ranking and the tradeoffs, or read how we score.

If you sign up through links on this page, vsref may earn a commission; programs exist on both sides of most comparisons, and commissions never change verdicts. How we make money

Reviewed by vsref Editorialfacts verified Jul 24, 2026Methodology →

1Moonshine logoMoonshineWINNER
On-device streaming STT for live voice interfaces, from tiny edge models to Whisper Large v3-beating accuracy2 of 2 points · 1 matchup
See pricingWebsite →
Open-weights, GPU-accelerated self-hosted STT stack1.5 of 4 points · 2 matchups
Developer-first realtime STT API for voice agents and transcription at scale8 of 32 points · 16 matchups
See pricingTry Deepgram

What matters for dictation

Weighted attribute comparison for Dictation
FactMoonshineNVIDIA Parakeet / RivaDeepgramAssemblyAIElevenLabs Scribe
Platforms×5n/an/an/an/an/a
Local vs cloud processing×4n/an/an/an/an/a
AI formatting / edit features×4n/an/an/an/an/a
App pricing (one-time vs subscription)×3n/an/an/an/an/a
App integrations×3n/an/an/an/an/a
Swipe → to see every tool column.
×5 Platforms: An app that does not run on your OS is out, whatever else it does.×4 Local vs cloud processing: On-device transcription keeps your words off vendor servers and works offline.×4 AI formatting / edit features: Auto-punctuation and cleanup are what separate dictation from raw transcription.×3 App pricing (one-time vs subscription): One-time licenses and subscriptions price out very differently over a year of use.×3 App integrations: System-wide dictation beats an app you have to copy-paste out of.

The ranking, tool by tool

1Moonshine logoMoonshineWINNER
For voice typing on your own computer, local processing is the dominant concern because users want audio to stay on device.
See pricingWebsite →

For voice typing on your own computer, local processing is the dominant concern because users want audio to stay on device. Full Moonshine vs OpenAI Whisper (API) verdict →

For voice typing on your own computer, local processing is the top priority after platform fit.

For voice typing on your own computer, local processing is the top priority after platform fit. Full NVIDIA Parakeet / Riva vs OpenAI Whisper (API) verdict →

Developer-first realtime STT API for voice agents and transcription at scale.
See pricingTry Deepgram

Developer-first realtime STT API for voice agents and transcription at scale. No won verdicts for this use case yet; it ranks on ties and near-misses.

Accuracy-led voice AI API for developers and voice agents.

Accuracy-led voice AI API for developers and voice agents. No won verdicts for this use case yet; it ranks on ties and near-misses.

Accuracy-led STT API from the leading AI audio company.

Accuracy-led STT API from the leading AI audio company. No won verdicts for this use case yet; it ranks on ties and near-misses.

EU-based real-time and batch STT API built on the Solaria models.
See pricingTry Gladia

EU-based real-time and batch STT API built on the Solaria models. No won verdicts for this use case yet; it ranks on ties and near-misses.

Hyperscaler STT API with Chirp foundation models and enterprise compliance.

Hyperscaler STT API with Chirp foundation models and enterprise compliance. No won verdicts for this use case yet; it ranks on ties and near-misses.

Low-cost EU-based transcription API with Apache-2.0 open-weight models.

Low-cost EU-based transcription API with Apache-2.0 open-weight models. No won verdicts for this use case yet; it ranks on ties and near-misses.

AWS-native STT API with deep AWS ecosystem integration and compliance coverage.

AWS-native STT API with deep AWS ecosystem integration and compliance coverage. No won verdicts for this use case yet; it ranks on ties and near-misses.

Enterprise-grade STT inside the Azure cloud ecosystem.

Enterprise-grade STT inside the Azure cloud ecosystem. No won verdicts for this use case yet; it ranks on ties and near-misses.

Transcription-heritage STT API with low per-hour pricing and open (non-commercial) Reverb models.
See pricingTry Rev AI

Transcription-heritage STT API with low per-hour pricing and open (non-commercial) Reverb models. No won verdicts for this use case yet; it ranks on ties and near-misses.

Ultra-low-cost multilingual STT + real-time translation API.
See pricingTry Soniox

Ultra-low-cost multilingual STT + real-time translation API. No won verdicts for this use case yet; it ranks on ties and near-misses.

Accuracy-first enterprise STT with flexible deployment (SaaS, container, on-prem).

Accuracy-first enterprise STT with flexible deployment (SaaS, container, on-prem). No won verdicts for this use case yet; it ranks on ties and near-misses.

Streaming STT for voice agents with native turn detection.

Streaming STT for voice agents with native turn detection. No won verdicts for this use case yet; it ranks on ties and near-misses.

Enterprise-grade open ASR for accurate batch transcription.

Enterprise-grade open ASR for accurate batch transcription. No won verdicts for this use case yet; it ranks on ties and near-misses.

Flagship GPT-4o based transcription API from OpenAI.

Flagship GPT-4o based transcription API from OpenAI. No won verdicts for this use case yet; it ranks on ties and near-misses.

Low-cost hosted STT API on the Grok stack.

Low-cost hosted STT API on the Grok stack. No won verdicts for this use case yet; it ranks on ties and near-misses.

Ultra-fast, low-cost hosted Whisper transcription API (no realtime streaming).

Ultra-fast, low-cost hosted Whisper transcription API (no realtime streaming). No won verdicts for this use case yet; it ranks on ties and near-misses.

Open-weights multilingual ASR models.
See pricingWebsite →

Open-weights multilingual ASR models. No won verdicts for this use case yet; it ranks on ties and near-misses.

Ultra-low-latency multilingual STT for voice agents.

Ultra-low-latency multilingual STT for voice agents. No won verdicts for this use case yet; it ranks on ties and near-misses.

Low-cost pay-as-you-go file transcription from OpenAI; no first-party streaming or diarization.

Low-cost pay-as-you-go file transcription from OpenAI; no first-party streaming or diarization. No won verdicts for this use case yet; it ranks on ties and near-misses.

More Speech-to-text APIs buyer guides