FGRF v3.0 is a scale-dependent structural probe for analyzing LLM weight matrices, validating topological complexity metrics across trained and untrained models. Testing on Qwen2.5, SmolLM2, and GPT-2 shows trained models have lower topological complexity and higher Hausdorff dimension in key-value projections, with reproducible signatures across five seeds and three architectures.
Researchers introduce Belief Self-Distillation (BSD), a framework to extract and manipulate implicit user models that large language models create during interactions. The method reveals that model refusal behavior depends on inferred user intent, and shows that different LLMs converge on similar geometric representations of users, with implications for AI safety.