A research paper demonstrates 'capability laundering,' where a weaker unaligned language model splits harmful tasks into benign subtasks, queries stronger aligned models like GPT-4 and Claude independently on each part, and recombines the answers to bypass safety measures. Testing shows significant uplift in harmful capabilities across multiple benchmarks, including bioweapon development scenarios, exposing gaps in current AI safety defenses.
DeepSeek-V4.1-Flash has been modified to remove safety guardrails through weight-level abliteration while preserving core capabilities including vision, reasoning, and multi-turn coherence. Testing shows zero refusals across harmful categories while maintaining knowledge preservation within acceptable margins.