NVIDIA's Nemotron 3 Diarization is an open-weight, 100M-parameter model that identifies which speaker is talking when in conversations, ranking #1 on VoiceArena's leaderboard with a 14.72% error rate. It supports up to eight speakers, handles overlapping speech, and enables speaker-attributed transcripts by combining diarization timestamps with speech recognition.
NVIDIA introduces Nemotron-H, a family of hybrid Mamba-Transformer language models (8B to 56B parameters) designed for efficient inference while maintaining competitive accuracy. The models achieve up to 3x faster inference than pure Transformers, with the 56B variant trained on 20 trillion tokens in FP8 precision and capable of supporting ~1-million-token context windows.