A Bitter Lesson for Data Filtering

5 pointsposted 10 hours ago
by nujan_dev

1 Comments

nujan_dev

10 hours ago

Large models don't just tolerate noisy or nominally "low-quality" web data but this paper suggests that they benefit from the distributional entropy in it. Over-curated datasets artificially compress the representation manifold, starving the model of the edge cases needed for robust generalization.