• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    User data is unlikely to add much to frontier AI’s math gains, one post argues

    The author calls for stronger disclosure norms around how AI companies train on user data, saying firms vary in how aggressively they use it and should explain their methods and intended capability improvements.

    SH
    AG
    ME
    6 Sources, ,

    TLDR

    One post argues that frontier models’ gains in areas like math come from scaling pretraining and reinforcement learning with verifiable rewards, with user-data training exceedingly unlikely to contribute much. It suggests user data is more likely used to find failure modes or situations that hired annotators struggle to recreate, and calls for clearer disclosure of training practices. A related comment questions whether frontier labs would want Chegg’s user data, pointing skeptically to homework-solution data from struggling undergraduates.

    Combined views

    41.9K

    6 Sources, first seen 20d ago

    Combined views

    41.9K

    6 Sources, first seen 20d ago

    360 likes
    20d ago
    first seen 20d ago
    360 likes
    11 comments
    189 saves
    147 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    11 comments
    189 saves
    147 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    6 Sources

    @menhguinsomeone once asked me why $CHGG couldn't just pivot to selling their user data to frontier labs you probably don't want to put homework solution data from struggling undergrads into a frontier model
    @sarahookr@figma @Lovable How do they do this. There are numerous ways. @johnschulman2 shared a few in this thread: Additionally there are clever synthetic data techniques that can generate distributional equivalent data while preserving privacy.
    @mervenoyannRT @johnschulman2: This isn't the most notable aspect of today's news, but on the user data issue, there are different kinds of *training o…
    @herbiebradley@sarahookr @figma @Lovable @johnschulman2 Under enterprise ZDR all of the methods John lists would seem to be forbidden, right? My understanding is for enterprise use today, FDEs are a big channel for figuring out what new RL envs to build, due to the ZDR.
    @animesh_gargwonder if robotics use cases would want similar data privacy. No data flows back.

    6 Sources

    @menhguinsomeone once asked me why $CHGG couldn't just pivot to selling their user data to frontier labs you probably don't want to put homework solution data from struggling undergrads into a frontier model
    @sarahookr@figma @Lovable How do they do this. There are numerous ways. @johnschulman2 shared a few in this thread: Additionally there are clever synthetic data techniques that can generate distributional equivalent data while preserving privacy.
    @mervenoyannRT @johnschulman2: This isn't the most notable aspect of today's news, but on the user data issue, there are different kinds of *training o…
    @herbiebradley@sarahookr @figma @Lovable @johnschulman2 Under enterprise ZDR all of the methods John lists would seem to be forbidden, right? My understanding is for enterprise use today, FDEs are a big channel for figuring out what new RL envs to build, due to the ZDR.
    @animesh_gargwonder if robotics use cases would want similar data privacy. No data flows back.