A CPU-based voice AI system that enables real-time speech-to-text and synthesis without requiring GPUs, deployable on cloud, on-premises, or edge devices. The pipeline includes noise suppression, turn-taking, transcription, and voice synthesis components that can run on standard processors.