# Kyutai STT Review

> Kyutai STT review: Streaming-first open STT for self-hosted voice agents. Verified pricing, features, and the strongest alternatives in Speech-to-text APIs.

![Kyutai STT logo](https://www.versusref.com/logos/kyutai-stt.png)

Streaming-first open STT for self-hosted voice agents

Among the 43 speech-to-text tools we track, Kyutai STT has the 42nd-widest language coverage.

See pricing · Open source · facts verified Jul 20, 2026

## What we know about Kyutai STT

Kyutai STT sits in the speech-to-text apis category, where it is streaming-first open STT for self-hosted voice agents. We keep this Kyutai STT profile grounded in primary sources, each fact dated to when we last confirmed it.

Kyutai STT does not publish a public per-minute rate we have been able to verify, so the price figure here stays blank until we can confirm one. We would rather show nothing than a guessed number; if you have a current Kyutai STT quote, it is the fastest way for us to close that gap.

On capabilities, Kyutai STT covers word-level timestamps, self-host / on-prem option, and websocket streaming api. Each of those is verified against Kyutai STT's own docs or dashboard, not marketing copy.

Among the 43 speech-to-text tools in our matrix, Kyutai STT leads with the 42nd-widest language coverage; the fact sheet below has the raw numbers behind that placement.

In total we track 14 verified facts for Kyutai STT today, each linking the primary source it came from so you can check our work - and vendor claims we have not measured ourselves are labeled as such on the Kyutai STT fact sheet below.

## Fact sheet

### Capabilities

| Fact | Value | Verified | Source |
| --- | --- | --- | --- |
| WER (vendor-claimed) | ~6.4 % WER | Jul 20 | [source](https://huggingface.co/kyutai/stt-2.6b-en) |
| Streaming latency (vendor-claimed) | ~500 ms | Jul 20 | [source](https://kyutai.org/stt) |
| Languages supported | 2 languages | Jul 20 | [source](https://github.com/kyutai-labs/delayed-streams-modeling) |
| Word-level timestamps | ✓  Yes | Jul 20 | [source](https://github.com/kyutai-labs/delayed-streams-modeling) |

### Compliance & trust

| Fact | Value | Verified | Source |
| --- | --- | --- | --- |
| Self-host / on-prem option | ✓  Yes | Jul 20 | [source](https://github.com/kyutai-labs/delayed-streams-modeling) |
| Model weights license | CC-BY-4.0 | Jul 20 | [source](https://huggingface.co/kyutai/stt-1b-en_fr) |

### Build experience

| Fact | Value | Verified | Source |
| --- | --- | --- | --- |
| Official SDKs | Python (PyTorch), Rust, MLX | Jul 20 | [source](https://github.com/kyutai-labs/delayed-streams-modeling) |
| Max file size / duration | Handles audio up to 2 hours per sequence | Jul 20 | [source](https://huggingface.co/kyutai/stt-1b-en_fr) |
| Supported audio formats | 24 kHz mono audio | Jul 20 | [source](https://huggingface.co/kyutai/stt-1b-en_fr) |
| Websocket streaming API | ✓  Yes | Jul 20 | [source](https://github.com/kyutai-labs/delayed-streams-modeling) |

### Commercial

| Fact | Value | Verified | Source |
| --- | --- | --- | --- |
| Model size (parameters) | 1B (en/fr), 2.6B (en) | Jul 20 | [source](https://github.com/kyutai-labs/delayed-streams-modeling) |
| Hardware to self-host | GPU for serving; MLX on-device on Apple silicon | Jul 20 | [source](https://kyutai.org/stt) |
| Project maintenance status | Active - latest commit 2026-01-26 | Jul 20 | [source](https://github.com/kyutai-labs/delayed-streams-modeling/commits/main) |
| GitHub stars | 2,980 stars | Jul 20 | [source](https://github.com/kyutai-labs/delayed-streams-modeling) |

Source: https://www.versusref.com/stt/tools/kyutai-stt/
