DeepSeek AI released DeepSeek-V4.1-Flash, a multimodal mixture-of-experts model with 1 million token context window, featuring 552B main parameters plus 196B Engram parameters. The model uses FP4 KV cache and cross-layer attention reuse to reduce memory consumption to 890 bytes per token, approximately 1/4 of DeepSeek-V4-Flash and 1/437 of DeepSeek-V1.