A Bitter Lesson for Data Filtering
Exploring the necessity of data filtering in large model pretraining. Low-quality data may provide unexpected benefits.
This study examines data filtering in large model pretraining through scaling experiments in high compute, data-scarce environments. Contrary to the belief that only high-quality data is essential, our findings indicate that with sufficient computational resources, the optimal approach may be to forgo data filtering altogether. Well-trained large parameter models not only handle low-quality and distractor data but can actually benefit from what is traditionally viewed as 'poor' data.
This synthesis was produced from its source by AI; there is no human editor or manual review step. How we work