DeepSeek released the V4.1 Flash model, a 552B parameter MoE model with native multimodal vision capabilities that outperforms V4 Pro across benchmarks while significantly reducing API pricing. The model features a new Causal-Encoder-Decoder architecture with dramatically reduced KV Cache requirements and improved inference speed, with V4 Pro being phased out in favor of the new model.
DeepSeek released V4.1 Flash on September 10, 2026, and postponed discontinuation of V4 Pro until September 14, 2026, when Pro requests will be routed to Flash at Flash pricing. V4.1 Flash has surpassed V4 Pro across performance, cost, speed, and task completion metrics.