NVIDIA's Nemotron 3 Diarization is an open-weight, 100M-parameter model that identifies which speaker is talking when in conversations, ranking #1 on VoiceArena's leaderboard with a 14.72% error rate. It supports up to eight speakers, handles overlapping speech, and enables speaker-attributed transcripts by combining diarization timestamps with speech recognition.
NobodyWho is an inference engine for running large language models locally and offline across multiple platforms including Android, iOS, and desktop. It supports various LLMs in GGUF format with features like multimodal input, text-to-speech, speech-to-text, and GPU acceleration via Vulkan or Metal. The tool provides SDKs for Kotlin, Swift, Python, Flutter, React Native, Expo, and Godot.
ScribeToAny is a free transcription service that converts audio and video files into text across 98 languages, using advanced Whisper technology to handle various accents and noisy backgrounds while supporting multiple export formats.
Synthesia launched its Interactive Avatar API, enabling developers to embed real-time, photorealistic avatars into web and app products. The API uses a bring-your-own-stack approach, allowing integration with custom LLMs, speech-to-text, and text-to-speech providers while Synthesia handles avatar rendering. The service supports 240+ avatars in 160+ languages and is now available for Enterprise customers in applications like customer support, HR onboarding, and interactive kiosks.
Fluentry is a local voice-to-text dictation tool for Linux and Wayland that runs speech recognition models on your machine without uploading data. It supports dozens of languages, integrates with PipeWire and system keyboards, and includes features like custom dictionary management, filler-word removal, and optional local grammar correction through language models.
Wispr Advanced Interfaces Lab introduced Canto, a speech recognition model designed for real-world dictation conditions. Tested on real Wispr Flow usage data, Canto achieved the lowest word error rate among comparable models from Google, OpenAI, AssemblyAI, and Deepgram, performing particularly well on low-volume and short utterances despite ranking second to Gemini 3.1 Pro on overall challenge audio.