Paper on Repetition Mismatch in Data Mixtures Announced
Kevin Zhou shares findings from work on dataset repetitions in limited data settings.
TLDR
Kevin Zhou announced on X that the paper Repetition Mismatch: Why Data Mixture Experiments Don’t Scale and How to Fix Them will appear in Budapest. The research looks at how repeating datasets influences the best way to mix data when supplies are limited. Zhou works as a machine learning scientist with AbridgeHQ, Imperial College, and Cornell. The post features the paper title on a white background and links to further discussion in a thread.
Combined views
7.3K
2 Sources, first seen 26d ago