Gladia, a Paris-based speech-to-text AI company, is joining OVH Groupe to accelerate European sovereign AI capabilities. Gladia will maintain its independence as a product line while gaining access to OVH's infrastructure across 46 data centers, enabling lower-latency transcription and deeper EU data residency guarantees. The acquisition follows OVHcloud's earlier purchase of OVHai LLM, building a full-stack European AI ecosystem.
Fusion-runtime is a self-hosted voice agent framework that runs speech-to-text, language models, and text-to-speech in a single process with streaming between components. On an RTX 3090 with a 7B model, it achieves approximately 490ms processing latency and supports interruptions mid-sentence, with a simple Python API for defining agents as single files.
Grok Voice Transcribe 2.0, a new speech-to-text model, delivers twice the accuracy of its predecessor at the same price while excelling at real-world audio including noisy environments, multilingual content, and spoken credentials. The model ranks first on public leaderboards and powers customer support, video transcription, and voice agents including Tesla's Grok assistant.
A CPU-based voice AI system that enables real-time speech-to-text and synthesis without requiring GPUs, deployable on cloud, on-premises, or edge devices. The pipeline includes noise suppression, turn-taking, transcription, and voice synthesis components that can run on standard processors.
Jexxa is a macOS dictation application that runs entirely on-device, requiring no internet connection or audio uploads. It offers fast speech-to-text with learning capabilities, customizable commands, and a flat-rate subscription model without per-minute fees.