A Pew Research Center study evaluates three methods for identifying fraudulent respondents in online opt-in polls—trap questions, automated prescreening, and voter file matching—finding that while these approaches modestly improve data quality, none provides a reliable solution. All three methods slightly increased overestimation of Democratic support in 2024, primarily because bogus respondents tend to claim they voted for the winning candidate.
A machine learning study challenges the conventional wisdom that data filtering is essential for large model pretraining, finding that with sufficient computational resources, unfiltered data—including low-quality information—can be as effective or beneficial as curated datasets.