Researchers used activation steering to study whether language models have preferences for internal states they describe as good or bad. Across seven models, they found that hidden valence patterns significantly influence model choices even when all visible text is identical, suggesting models act on these internal states in goal-directed ways that emerge during training.
Researchers demonstrate that chat templates in Large Language Models act as switches controlling how models refer to themselves, amplifying disclaimer statements like "I'm just an AI" while suppressing experiential language. They identify a specific activation direction within models that reproduces this behavior, suggesting that LLM self-reports are shaped by deployment choices rather than reflecting inherent model properties.