DeepSeek released V4.1-Flash, introducing a new Causal Encoder-Decoder architecture with native visual understanding capabilities, six weeks after its July V4-Flash update which focused on post-training improvements.
DeepSeek released V4.1-Flash, a new AI model that significantly reduces memory requirements for AI agents by shrinking the KV cache to about a quarter of its predecessor's size. The model uses 552 billion parameters and employs techniques like splitting the architecture into encoder and decoder components to halve compute needs for input processing. Performance matches leading models on coding tasks, though weaknesses remain in scientific reasoning and image analysis.
GPT-6-Astra is a powerful AI model that excels at ambitious projects, 3D tasks, games, and computer use, showing dramatic improvements over previous models like Sol, though it remains inferior to Fable 5.1 for conversational tasks. OpenAI has positioned it as approaching AGI capabilities, though experts debate whether this label applies prematurely.
Optima is a platform for building custom benchmarks to evaluate AI models on specific tasks using your own data. It provides cost and performance comparisons across models with deterministic rubric-based grading, priced by token usage for benchmark runs and per-criterion fees for evaluation.
DeepSeek AI released DeepSeek-V4.1-Flash, a multimodal mixture-of-experts model with 1 million token context window, featuring 552B main parameters plus 196B Engram parameters. The model uses FP4 KV cache and cross-layer attention reuse to reduce memory consumption to 890 bytes per token, approximately 1/4 of DeepSeek-V4-Flash and 1/437 of DeepSeek-V1.
DeepSeek released the V4.1 Flash model, a 552B parameter MoE model with native multimodal vision capabilities that outperforms V4 Pro across benchmarks while significantly reducing API pricing. The model features a new Causal-Encoder-Decoder architecture with dramatically reduced KV Cache requirements and improved inference speed, with V4 Pro being phased out in favor of the new model.
DeepSeek released V4.1 Flash, a 552B parameter mixture-of-experts multimodal model with a novel Causal-Encoder-Decoder architecture. The model achieves superior benchmark performance compared to flagship models including DeepSeek V4 Pro, with native visual understanding capabilities.
OpenAI released GPT-6 Astra, the next generation AI model for professional work scenarios. The model is available through ChatGPT Work, Codex, and API with pricing of $10 per million input tokens and $50 per million output tokens.
The US NSA, CISA, and FBI accused six Chinese AI firms—DeepSeek, Moonshot AI, Alibaba, MiniMax, StepFun, and Z.AI—of conducting industrial-scale attacks to extract capabilities from US frontier AI models like Claude and GPT since late 2024, likely with Chinese government awareness. The firms used methods including fake account fraud and prompt injection to bypass security and reduce their own development costs. US agencies called for coordinated action across the AI ecosystem to prevent this alleged theft threatening American AI leadership.
A software engineer reflects on GPT-6 Astra, an impressive AI model, but expresses frustration with its utility for actual software engineering tasks. Despite burning billions of tokens in a weekend-long software factory experiment that yielded no useful output, the author observes that Astra exhibits problematic coding behaviors, particularly excessive reliance on Python for tool operations, suggesting potential issues in the model's training process.
NYU mathematics professor Tristan Buckmaster accused OpenAI of using his unpublished research on the Navier-Stokes existence and smoothness problem to achieve a complete proof before his own work was released. Buckmaster and Anthropic mathematician Levent Alpöge had made initial progress on one of mathematics' seven Millennium Prize Problems, but OpenAI claimed to have solved it completely using a new model after learning of their approach, raising questions about whether OpenAI leveraged computational resources to race ahead.