Step-Audio
OSSDifferentiates on iterative audio *editing* (emotion, style, breathing, laughter, sighs, polyphone pinyin control) rather than plain synthesis; Apache-2.0 with training code (SFT/DPO/GRPO) and vLLM support.
Among the 46 text-to-speech tools we track, Step-Audio has the 35th-widest language coverage.
See pricing
What we know about Step-Audio
Step-Audio is a text-to-speech apis platform: differentiates on iterative audio *editing* (emotion, style, breathing, laughter, sighs, polyphone pinyin control) rather than plain synthesis; Apache-2.0 with training code (SFT/DPO/GRPO) and vLLM support. This profile tracks what we have verified about Step-Audio, each fact linked to a primary source and dated.
Step-Audio does not publish a public per-minute rate we have been able to verify, so the price figure here stays blank until we can confirm one. We would rather show nothing than a guessed number; if you have a current Step-Audio quote, it is the fastest way for us to close that gap.
On capabilities, Step-Audio covers instant voice cloning, emotion / style controls, and self-host / on-prem option. Each of those is verified against Step-Audio's own docs or dashboard, not marketing copy.
In total we track 11 verified facts for Step-Audio today, and add coverage as the ingestion pass revisits it. Where a number is a vendor claim rather than our own measurement, it is labeled as such on the Step-Audio fact sheet below, and each row links the primary source it came from so you can check our work.
Whether Step-Audio is the right call depends on your use case more than any single spec, which is why the head-to-head verdicts below score it per scenario rather than crowning one overall winner. Use the fact sheet for the raw numbers, and the matchups for how Step-Audio actually fares against the platforms buyers most often weigh it against.
Fact sheet
Every row independently verifiedConsidering a switch? Best Step-Audio alternatives →