Qwen-Audio-3.1-TTS is a production-oriented speech synthesis system combining a low-frame-rate tokenizer with progressive training to achieve state-of-the-art performance across content consistency, speaker similarity, prosody, and audio quality. It supports 16 languages and 20 Chinese dialects with fine-grained controllability through natural-language instructions and inline tags, enabling robust synthesis up to 3 minutes including handling of noisy reference speech.