DeepSeek-AI released DeepSeek-V4.1-Flash, a 552B mixture-of-experts multimodal model focused on advancing KV cache compression techniques to improve efficiency.
DeepSeek released V4.1-Flash, introducing a new Causal Encoder-Decoder architecture with native visual understanding capabilities, six weeks after its July V4-Flash update which focused on post-training improvements.
DeepSeek released V4.1-Flash, a new AI model that significantly reduces memory requirements for AI agents by shrinking the KV cache to about a quarter of its predecessor's size. The model uses 552 billion parameters and employs techniques like splitting the architecture into encoder and decoder components to halve compute needs for input processing. Performance matches leading models on coding tasks, though weaknesses remain in scientific reasoning and image analysis.
Anthropic reported coordinated distillation attacks by China-based AI companies Alibaba, Moonshot AI, and DeepSeek targeting Claude's capabilities. The company observed nearly 200 million exchanges across five campaigns designed to extract chain-of-thought reasoning to train competing models, with Alibaba's effort accounting for 151 million exchanges between May and July 2026.