• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Paper on Repetition Mismatch in Data Mixtures Announced

    Kevin Zhou shares findings from work on dataset repetitions in limited data settings.

    KC
    KZ
    2 Sources, 26d ago, first seen 26d ago

    TLDR

    Kevin Zhou announced on X that the paper Repetition Mismatch: Why Data Mixture Experiments Don’t Scale and How to Fix Them will appear in Budapest. The research looks at how repeating datasets influences the best way to mix data when supplies are limited. Zhou works as a machine learning scientist with AbridgeHQ, Imperial College, and Cornell. The post features the paper title on a white background and links to further discussion in a thread.

    Combined views

    7.3K

    2 Sources, first seen 26d ago

    Combined views

    7.3K

    2 Sources, first seen 26d ago

    63 likes
    63 likes
    2 comments
    39 saves
    11 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    2 comments
    39 saves
    11 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    2 Sources

    @kevinzhou497Our paper “Repetition Mismatch: Why Data Mixture Experiments Don’t Scale and How to Fix Them” will be at #EMNLP2026 in Budapest! In it, we investigate the role of dataset repetitions on optimal data mixing in data constrained settings 🧵⬇️
    @kroscooRepeat-aware is like muP for data mixtures because it lets you tune at a small scale and then transfer to your final intended run. Great work by @kevinzhou497 !

    2 Sources

    @kevinzhou497Our paper “Repetition Mismatch: Why Data Mixture Experiments Don’t Scale and How to Fix Them” will be at #EMNLP2026 in Budapest! In it, we investigate the role of dataset repetitions on optimal data mixing in data constrained settings 🧵⬇️
    @kroscooRepeat-aware is like muP for data mixtures because it lets you tune at a small scale and then transfer to your final intended run. Great work by @kevinzhou497 !