A researcher tested whether modern LLMs contain internal models of other LLMs by having Qwen complete text started by GPT-2, comparing whether Qwen's continuations resembled GPT-2's own continuations more than Qwen's natural output. The experiment used headlines from September 2026 (outside training cutoffs) and varied generation lengths to measure textual overlap between conditions.