A Bitter Lesson for Data Filtering
By Christopher Mohri, John Duchi, Tatsunori Hashimoto
Presents scaling studies showing that with sufficient compute in data-scarce regimes, the best data filter is no data filter—large models benefit from nominally 'poor' data. This challenges the common belief that aggressive data filtering is essential for pretraining.