Xing4.0-29B-A4B is a 29-billion-parameter mixture-of-experts language model released by China Telecom, trained entirely on Huawei's Ascend NPU hardware rather than Nvidia GPUs. The model activates only 4 billion parameters per token, supports up to 512K context windows, and is optimized for agentic tasks including tool calling, coding, and multi-step planning. It integrates with deployment frameworks like vLLM and agent platforms including Claude Code.
China's annual token consumption is projected to reach 100 quadrillion in 2026 and exceed 35 quintillion by 2030, driven by rapid deployment of AI agents across industries. The shift focuses on commercializing AI agents rather than competing on large models, with inference computing expected to dominate demand by 2029. Chinese telecom and mobile companies are preparing infrastructure to handle the massive growth in token consumption and computing power needs.