vsref
LMNT logo

LMNT Review

Speed-and-affordability challenger: a single fast model (Blizzard 2) with streaming-first APIs, unlimited voice clones on every tier, and flat subscription pricing with per-1K overage.

Among the 46 text-to-speech tools we track, LMNT has the 15th-widest language coverage - a fit for multilingual and localization projects.

From $10/mo

Facts verified Jul 20, 2026Try LMNT →

If you sign up through links on this page, vsref may earn a commission; programs exist on both sides of most comparisons, and commissions never change verdicts. How we make money

What we know about LMNT

LMNT sits in the text-to-speech apis category, where it is speed-and-affordability challenger: a single fast model (Blizzard 2) with streaming-first APIs, unlimited voice clones on every tier, and flat subscription pricing with per-1K overage. We keep this LMNT profile grounded in primary sources, each fact dated to when we last confirmed it.

On pricing, LMNT starts at $10 per month for its entry tier. That is the sticker rate: real production cost usually runs higher once you add a language model, a voice provider, and telephony minutes.

On capabilities, LMNT covers streaming audio output, realtime websocket api, instant voice cloning, emotion / style controls, and word-level timestamps. Each of those is verified against LMNT's own docs or dashboard, not marketing copy.

Placed against the 46 text-to-speech tools we track, LMNT's strongest showing is the 15th-widest language coverage, while it trails at 9th on latency - a spread worth weighing against your own priorities.

In total we track 17 verified facts for LMNT today, each linking the primary source it came from so you can check our work - and vendor claims we have not measured ourselves are labeled as such on the LMNT fact sheet below.

Reviewed by vsref Editorialfacts verified Jul 20, 2026Methodology →

LMNT pricing

Cheapest paid plan $10/mo with 200,000 characters included, verified Jul 20, 2026 (source).

LMNT cost at monthly volume tiers
Monthly volume$ / 1M chars$ / audio-minMonthly bill
200K chars/moHobby project$50$0.0475$10
2M chars/moProduct feature$50$0.0475$100
20M chars/moAt scale$50$0.0475$1,000

Effective rates from published pricing; subscription plans resolve to plan fee plus overage. ~950 characters ≈ 1 audio minute. How we compute costs →

Fact sheet

Pricing
Pricing facts
Pricing modelhybridJul 20
Cheapest paid plan$10Jul 20
Characters included on cheapest plan200,000 chars/moJul 20
Free tier quota15K charactersJul 20
Overage rate past plan quota50 $/1M charsJul 20
Capabilities
Capabilities facts
TTFB latency (vendor-claimed)~150 msJul 20
Streaming audio output✓ YesJul 20
Realtime websocket API✓ YesJul 20
Instant voice cloning✓ YesJul 20
Minimum audio for voice cloning5-10 secondsJul 20
Languages supported~31 languagesJul 20
Emotion / style controls✓ YesJul 20
Word-level timestamps✓ YesJul 20
Build experience
Build experience facts
Official SDKsPython, TypeScript/Node, GoJul 20
Concurrency on base planNo concurrency or rate limits (paid plans)Jul 20
Output formatsaac, mp3, ulaw, wav, webm, pcm_s16le, pcm_f32leJul 20
Max input per request5000Jul 20

Considering a switch? Best LMNT alternatives →

STT in this stack

Text-to-speech is half of a voice pipeline: the other half is the speech-to-text that listens. Compare transcription engines on accuracy, streaming latency, and per-minute price: Deepgram, AssemblyAI, GPT-4o Transcribe, or the full speech-to-text comparison →

Voice agents in this stack

The engine is one layer: a voice agent speaks through its text-to-speech engine, but orchestration, telephony, and turn-taking come from the agent platform. Compare the platforms builders pair LMNT with: Pipecat, OpenAI Realtime API, Twilio ConversationRelay, or the full voice-agent comparison →

Distribute it

Most LMNT voiceover ends up in short-form video, and publishing that video across TikTok, YouTube, and Instagram is a scheduling problem with real per-channel pricing. Compare the schedulers creators actually run: Buffer, Postiz, Mixpost, or the social media scheduling platforms compared →

Avatar video in this stack

A cloned or bring-your-own LMNT voice does not have to stay audio-only: AI avatar video platforms lip-sync it onto a talking avatar for finished video. Compare the platforms: HeyGen, Synthesia, Hedra, or the full avatar-video comparison →

LMNT head-to-head

LMNT vs Cartesia →won 1 · lost 5 · tied 0
AudiobookswonContent CreatorslostDeveloperslostDubbinglostSelf-HostedlostVoice Agentslost