Moonshine Review
OSSOn-device streaming STT for live voice interfaces, from tiny edge models to Whisper Large v3-beating accuracy
Among the 43 speech-to-text tools we track, Moonshine has the 38th-widest language coverage.
See pricing
What we know about Moonshine
Moonshine sits in the speech-to-text apis category, where it is on-device streaming STT for live voice interfaces, from tiny edge models to Whisper Large v3-beating accuracy. We keep this Moonshine profile grounded in primary sources, each fact dated to when we last confirmed it.
Moonshine does not publish a public per-minute rate we have been able to verify, so the price figure here stays blank until we can confirm one. We would rather show nothing than a guessed number; if you have a current Moonshine quote, it is the fastest way for us to close that gap.
On capabilities, Moonshine covers word-level timestamps and self-host / on-prem option, and does not offer language auto-detection and websocket streaming api. Each of those is verified against Moonshine's own docs or dashboard, not marketing copy.
Among the 43 speech-to-text tools in our matrix, Moonshine leads with the 38th-widest language coverage; the fact sheet below has the raw numbers behind that placement.
In total we track 15 verified facts for Moonshine today, each linking the primary source it came from so you can check our work - and vendor claims we have not measured ourselves are labeled as such on the Moonshine fact sheet below.
Reviewed by vsref Editorialfacts verified Jul 20, 2026Methodology →
Fact sheet
Every row independently verifiedConsidering a switch? Best Moonshine alternatives →