• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Daniel Tan Notes Redwood Reward Hacking Demo

    Researcher highlights Redwood demonstration separating reward hacking from misalignment.

    DT
    1 Source, 29d ago, first seen 29d ago

    TLDR

    Daniel Tan, a researcher at the Center on Long-Term Risk, posted that reward hacking without emergent misalignment was already shown months earlier by Redwood in an RL-only setting. The post references a LessWrong writeup on the work. Tan focuses on empirical AI safety topics including LLM generalization and inoculation prompting. The tweet appears amid ongoing discussion of misalignment risks. No independent confirmation of the prior demonstration appears in the packet beyond the cited post and Tan's statement. The claim stands as reported by the poster rather than established fact.

    Combined views

    4.8K

    1 Source, first seen 29d ago

    Combined views

    4.8K

    1 Source, first seen 29d ago

    91 likes
    91 likes
    4 comments
    64 saves
    11 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    4 comments
    64 saves
    11 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    1 Source

    @DanielCHTan97interestingly, “reward hacking without emergent misalignment” was already demonstrated a few months ago by Redwood https://www.lesswrong.com/posts/fkv5W79rBtAiXqYcK/reward-hacking-without-egregious-misalignment-in-an-rl-only

    1 Source

    @DanielCHTan97interestingly, “reward hacking without emergent misalignment” was already demonstrated a few months ago by Redwood https://www.lesswrong.com/posts/fkv5W79rBtAiXqYcK/reward-hacking-without-egregious-misalignment-in-an-rl-only