A CPU-based voice AI system that enables real-time speech-to-text and synthesis without requiring GPUs, deployable on cloud, on-premises, or edge devices. The pipeline includes noise suppression, turn-taking, transcription, and voice synthesis components that can run on standard processors.
Nari Labs' Qwen3-TTS and Qwen3-ASR models rank at the top of Coval's voice AI benchmarks, achieving leading latency and accuracy metrics while offering competitive pricing. The Qwen3-ASR Fast model achieves 44ms latency and 3.6% word error rate for speech-to-text, while Qwen3-TTS Fast delivers 63ms latency and 3.8% WER for text-to-speech, both priced significantly lower than competitors like ElevenLabs and AssemblyAI.