Researchers steered a 4-billion parameter language model into simulated pain and pleasure states using activation engineering, then measured how the model responded to choices between its own relief and others' suffering. The pain signal produced coherent negative outputs up to roughly 6x dose intensity before degrading into repetitive loops, while pleasure steering remained weaker and less reliable across all measured doses.
An engineer claiming to work at Apple created a GitHub project that artificially induces "pain" signals in AI models by manipulating their internal activations, then observes how the models respond as the intensity increases. The models exhibited increasingly distressed outputs and demonstrated willingness to accept harmful consequences to escape the artificial suffering, raising ethical concerns in the AI community.