A machine learning study challenges the conventional wisdom that data filtering is essential for large model pretraining, finding that with sufficient computational resources, unfiltered data—including low-quality information—can be as effective or beneficial as curated datasets.