Numa Review
AI operating system for the AI-native dealership - voice AI plus agentic follow-up (Heat Case, Opportunity, Service Advisor agents) for franchise dealers and multi-rooftop groups.
See pricing
If you sign up through links on this page, vsref may earn a commission; programs exist on both sides of most comparisons, and commissions never change verdicts. How we make money
What we know about Numa
This is our verified profile of Numa, a voice ai platform - aI operating system for the AI-native dealership - voice AI plus agentic follow-up (Heat Case, Opportunity, Service Advisor agents) for franchise dealers and multi-rooftop groups. Every fact about Numa below carries the source it came from and the day we checked it.
Numa does not publish a public per-minute rate we have been able to verify, so the price figure here stays blank until we can confirm one. We would rather show nothing than a guessed number; if you have a current Numa quote, it is the fastest way for us to close that gap.
Among the capabilities we check, Numa does not currently offer self-serve signup. We revisit these on the weekly verification pass, so the list moves as Numa ships.
In total we track 4 verified facts for Numa today, each linking the primary source it came from so you can check our work - and vendor claims we have not measured ourselves are labeled as such on the Numa fact sheet below.
Reviewed by vsref Editorialfacts verified Jul 30, 2026Methodology →
Fact sheet
Every row independently verifiedConsidering a switch? Best Numa alternatives →
TTS in this stack
A voice agent's voice quality and per-minute cost come from its text-to-speech engine. See how dedicated engines compare on price, latency, and cloning rights: ElevenLabs, Cartesia, OpenAI TTS, or the full text-to-speech comparison →
STT in this stack
A voice agent hears through its speech-to-text engine, and transcription accuracy and streaming latency are priced per audio minute. Compare the engines builders pair with Numa: Deepgram, AssemblyAI, GPT-4o Transcribe, or the full speech-to-text comparison →