The author investigates why their 163M-parameter GPT-2 models underperform OpenAI's original GPT-2 despite using similar architecture. After ruling out factors like weight-tying and dropout handling, they identify training data quality as the likely culprit: OpenAI's WebText dataset, curated from highly-upvoted Reddit links, appears superior to FineWeb, the general web-scraping dataset the author used.