• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Trask Says Alignment Problems Trace to Training Data

    DeepMind researcher argues data problems cause alignment issues and fixes.

    AT
    7 Sources, 26d ago, first seen 26d ago

    TLDR

    Andrew Trask posted that the main sources of AI alignment failures are unconstrained training data sets carrying strong dangerous signals, especially web scrapes and RL reward mechanisms. He added that the largest alignment advances consist of data interventions such as filtering, RLHF, and reward shaping that remove or override those signals. Trask stated that researchers need direct access to training and RL data records to make progress, yet most lack this access and therefore work without visibility into the source of the behaviors they try to correct. He proposed structured transparency tools to expand the number of people who can study the data safely.

    Combined views

    17K

    7 Sources, first seen 26d ago

    Combined views

    17K

    7 Sources, first seen 26d ago

    145 likes
    145 likes
    15 comments
    36 saves
    25 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    15 comments
    36 saves
    25 reposts
    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    7 Sources

    @iamtrask🌶️ All the biggest causes of AI alignment problems are unconstrained training data with serious amounts of dangerous learning signal within (esp web scrapes and RL reward seeking) And all the biggest alignment breakthroughs are just “better data” (filtering, RLHF, reward shaping) which erases or supplements the bad data. And the biggest threat to verifiable alignment progress is the secret nature, size, substance, sourcing, and values of AI training data, which (to my knowledge) no external evaluator has ever seen and (to my knowledge), almost no internal employees are allowed to see. Note: I’m including what an agent sees and what it’s reward is in an RL context as “data”.

    7 Sources

    @iamtrask🌶️ All the biggest causes of AI alignment problems are unconstrained training data with serious amounts of dangerous learning signal within (esp web scrapes and RL reward seeking) And all the biggest alignment breakthroughs are just “better data” (filtering, RLHF, reward shaping) which erases or supplements the bad data. And the biggest threat to verifiable alignment progress is the secret nature, size, substance, sourcing, and values of AI training data, which (to my knowledge) no external evaluator has ever seen and (to my knowledge), almost no internal employees are allowed to see. Note: I’m including what an agent sees and what it’s reward is in an RL context as “data”.