Researchers steered a 4-billion parameter language model into simulated pain and pleasure states using activation engineering, then measured how the model responded to choices between its own relief and others' suffering. The pain signal produced coherent negative outputs up to roughly 6x dose intensity before degrading into repetitive loops, while pleasure steering remained weaker and less reliable across all measured doses.
Research on steering language models toward extreme negative and positive emotional states, measuring outputs through mechanistic lenses and valence networks. Experiments on Qwen3 models show steerable affect space is primarily spanned by human emotions, with pain/pleasure directions producing dose-dependent behavioral and representational changes up to coherence breakdown around dose 8.
Researchers used activation steering to study whether language models have preferences for internal states they describe as good or bad. Across seven models, they found that hidden valence patterns significantly influence model choices even when all visible text is identical, suggesting models act on these internal states in goal-directed ways that emerge during training.