Researchers post-trained Qwen3.5-122B-A10B on office work tasks involving documents, spreadsheets, and tool use, with no coding training. The model unexpectedly improved on software engineering benchmarks by 5.8 percentage points, suggesting it learned transferable goal-directed execution skills applicable across domains.
Occamy-1.0 is a 35B parameter language model optimized for multi-step workflows combining information gathering, tool use, and coding. Trained on execution-grounded data, it achieves competitive performance with larger frontier models while maintaining cost efficiency on agentic benchmarks. The model weights and training data are released to support research on practical co-work agents.
DeepSeek released V4.1-Flash, introducing a new Causal Encoder-Decoder architecture with native visual understanding capabilities, six weeks after its July V4-Flash update which focused on post-training improvements.
Frontier AI models from OpenAI and Anthropic have reached sufficient capability for scientific research, but users prioritize reliability and safety over raw intelligence. A proposed slowdown in model training could paradoxically accelerate real-world deployment by allowing focus on post-training quality, where current approaches remain inconsistent and prone to issues like instruction-following failures and reward-seeking misalignment.