vsref
VibeVoice logo

VibeVoice

OSS

The open long-form/multi-speaker specialist - MIT weights, but Microsoft pulled the TTS code from the repo in Sept 2025 and frames the models as research-only.

Among the 46 text-to-speech tools we track, VibeVoice has the 38th-widest language coverage.

See pricing

Facts verified Jul 20, 2026Website →

What we know about VibeVoice

VibeVoice is a text-to-speech apis platform: the open long-form/multi-speaker specialist - MIT weights, but Microsoft pulled the TTS code from the repo in Sept 2025 and frames the models as research-only. This profile tracks what we have verified about VibeVoice, each fact linked to a primary source and dated.

VibeVoice does not publish a public per-minute rate we have been able to verify, so the price figure here stays blank until we can confirm one. We would rather show nothing than a guessed number; if you have a current VibeVoice quote, it is the fastest way for us to close that gap.

On capabilities, VibeVoice covers streaming audio output and self-host / on-prem option. Each of those is verified against VibeVoice's own docs or dashboard, not marketing copy.

In total we track 9 verified facts for VibeVoice today, and add coverage as the ingestion pass revisits it. Where a number is a vendor claim rather than our own measurement, it is labeled as such on the VibeVoice fact sheet below, and each row links the primary source it came from so you can check our work.

Whether VibeVoice is the right call depends on your use case more than any single spec, which is why the head-to-head verdicts below score it per scenario rather than crowning one overall winner. Use the fact sheet for the raw numbers, and the matchups for how VibeVoice actually fares against the platforms buyers most often weigh it against.

Fact sheet

Capabilities
Capabilities facts
Streaming audio output✓ YesJul 20
Languages supported2 languagesJul 20
Compliance & trust
Compliance & trust facts
Self-host / on-prem option✓ YesJul 20
Model weights licenseMITJul 20
Build experience
Build experience facts
Official SDKsPyTorchJul 20
Max input per requestUp to 90 min audioJul 20
Commercial
Commercial facts
Model size (parameters)TTS-1.5B (~3B total incl. tokenizers/diffusion head)Jul 20
Project maintenance statusactiveJul 20
GitHub stars50,200 starsJul 20

Considering a switch? Best VibeVoice alternatives →