Qwopus3.8-27B-Flash-V2 is a post-trained language model based on Qwen3.8-27B, designed to reduce inefficient reasoning while maintaining problem-solving capability for agent workloads. The V2 update applies new reward functions and reinforcement-learning methods to improve inference speed and consistency without sacrificing task completion accuracy.
NVIDIA introduces Nemotron-H, a family of hybrid Mamba-Transformer language models (8B to 56B parameters) designed for efficient inference while maintaining competitive accuracy. The models achieve up to 3x faster inference than pure Transformers, with the 56B variant trained on 20 trillion tokens in FP8 precision and capable of supporting ~1-million-token context windows.