# Step-Audio Review

> Step-Audio review: Differentiates on iterative audio *editing* (emotion, style, breathing, laughter, sighs, polyphone pinyin control) rather than plain.

![Step-Audio logo](https://www.versusref.com/logos/step-audio.png)

Differentiates on iterative audio *editing* (emotion, style, breathing, laughter, sighs, polyphone pinyin control) rather than plain synthesis; Apache-2.0 with training code (SFT/DPO/GRPO) and vLLM support.

Among the 46 text-to-speech tools we track, Step-Audio has the 35th-widest language coverage.

See pricing · Open source · facts verified Jul 20, 2026

## What we know about Step-Audio

This is our verified profile of Step-Audio, a text-to-speech apis platform - differentiates on iterative audio *editing* (emotion, style, breathing, laughter, sighs, polyphone pinyin control) rather than plain synthesis; Apache-2.0 with training code (SFT/DPO/GRPO) and vLLM support. Every fact about Step-Audio below carries the source it came from and the day we checked it.

Step-Audio does not publish a public per-minute rate we have been able to verify, so the price figure here stays blank until we can confirm one. We would rather show nothing than a guessed number; if you have a current Step-Audio quote, it is the fastest way for us to close that gap.

On capabilities, Step-Audio covers instant voice cloning, emotion / style controls, and self-host / on-prem option. Each of those is verified against Step-Audio's own docs or dashboard, not marketing copy.

Placed against the 46 text-to-speech tools we track, Step-Audio's strongest showing is the 35th-widest language coverage - a spread worth weighing against your own priorities.

In total we track 11 verified facts for Step-Audio today, each linking the primary source it came from so you can check our work - and vendor claims we have not measured ourselves are labeled as such on the Step-Audio fact sheet below.

## Fact sheet

### Capabilities

| Fact | Value | Verified | Source |
| --- | --- | --- | --- |
| Instant voice cloning | ✓  Yes | Jul 20 | [source](https://github.com/stepfun-ai/Step-Audio-EditX) |
| Languages supported | 6 languages/dialects | Jul 20 | [source](https://github.com/stepfun-ai/Step-Audio-EditX) |
| Emotion / style controls | ✓  Yes | Jul 20 | [source](https://github.com/stepfun-ai/Step-Audio-EditX) |

### Compliance & trust

| Fact | Value | Verified | Source |
| --- | --- | --- | --- |
| Self-host / on-prem option | ✓  Yes | Jul 20 | [source](https://github.com/stepfun-ai/Step-Audio-EditX) |
| Model weights license | Apache-2.0 | Jul 20 | [source](https://github.com/stepfun-ai/Step-Audio-EditX) |

### Build experience

| Fact | Value | Verified | Source |
| --- | --- | --- | --- |
| Official SDKs | Python (Gradio demo, vLLM inference, training scripts) | Jul 20 | [source](https://github.com/stepfun-ai/Step-Audio-EditX) |

### Commercial

| Fact | Value | Verified | Source |
| --- | --- | --- | --- |
| Model size (parameters) | 3B | Jul 20 | [source](https://github.com/stepfun-ai/Step-Audio-EditX) |
| Hardware to self-host | ~12 GB GPU memory (16 GB recommended), CUDA | Jul 20 | [source](https://github.com/stepfun-ai/Step-Audio-EditX) |
| Hosted API available | Yes | Jul 20 | [source](https://github.com/stepfun-ai/Step-Audio-EditX) |
| Project maintenance status | active | Jul 20 | [source](https://github.com/stepfun-ai/Step-Audio-EditX) |
| GitHub stars | 951 stars | Jul 20 | [source](https://github.com/stepfun-ai/Step-Audio-EditX) |

Source: https://www.versusref.com/tts/tools/step-audio/
