Researchers introduce Belief Self-Distillation (BSD), a framework to extract and manipulate implicit user models that large language models create during interactions. The method reveals that model refusal behavior depends on inferred user intent, and shows that different LLMs converge on similar geometric representations of users, with implications for AI safety.