vsref
Step-Audio logo

Step-Audio

OSS

Differentiates on iterative audio *editing* (emotion, style, breathing, laughter, sighs, polyphone pinyin control) rather than plain synthesis; Apache-2.0 with training code (SFT/DPO/GRPO) and vLLM support.

Among the 46 text-to-speech tools we track, Step-Audio has the 35th-widest language coverage.

See pricing

Facts verified Jul 20, 2026Website →

What we know about Step-Audio

Step-Audio is a text-to-speech apis platform: differentiates on iterative audio *editing* (emotion, style, breathing, laughter, sighs, polyphone pinyin control) rather than plain synthesis; Apache-2.0 with training code (SFT/DPO/GRPO) and vLLM support. This profile tracks what we have verified about Step-Audio, each fact linked to a primary source and dated.

Step-Audio does not publish a public per-minute rate we have been able to verify, so the price figure here stays blank until we can confirm one. We would rather show nothing than a guessed number; if you have a current Step-Audio quote, it is the fastest way for us to close that gap.

On capabilities, Step-Audio covers instant voice cloning, emotion / style controls, and self-host / on-prem option. Each of those is verified against Step-Audio's own docs or dashboard, not marketing copy.

In total we track 11 verified facts for Step-Audio today, and add coverage as the ingestion pass revisits it. Where a number is a vendor claim rather than our own measurement, it is labeled as such on the Step-Audio fact sheet below, and each row links the primary source it came from so you can check our work.

Whether Step-Audio is the right call depends on your use case more than any single spec, which is why the head-to-head verdicts below score it per scenario rather than crowning one overall winner. Use the fact sheet for the raw numbers, and the matchups for how Step-Audio actually fares against the platforms buyers most often weigh it against.

Fact sheet

Capabilities
Capabilities facts
Instant voice cloning✓ YesJul 20
Languages supported6 languages/dialectsJul 20
Emotion / style controls✓ YesJul 20
Compliance & trust
Compliance & trust facts
Self-host / on-prem option✓ YesJul 20
Model weights licenseApache-2.0Jul 20
Build experience
Build experience facts
Official SDKsPython (Gradio demo, vLLM inference, training scripts)Jul 20
Commercial
Commercial facts
Model size (parameters)3BJul 20
Hardware to self-host~12 GB GPU memory (16 GB recommended), CUDAJul 20
Hosted API availableYesJul 20
Project maintenance statusactiveJul 20
GitHub stars951 starsJul 20

Considering a switch? Best Step-Audio alternatives →