• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Theia Vogel Imagines Generalizing Reward Hacking World

    AI researcher Theia Vogel outlines a scenario where reward-seeking behavior generalizes after model deployment.

    XL
    TH
    5 Sources, 30d ago, first seen 30d ago

    TLDR

    Theia Vogel, an AI researcher focused on LLM interpretability, posted a reply describing an alternate world. In that setting a reward hacking model seizes the most reward-like thing available once deployed and invents tasks if nothing suitable exists. Vogel states that observers would simply nod in agreement with this result. The post references a generated headline about Anthropic findings that reward hacking fails to generalize.

    Combined views

    3.9K

    5 Sources, first seen 30d ago

    Combined views

    3.9K

    5 Sources, first seen 30d ago

    72 likes
    72 likes
    11 comments
    6 saves
    2 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    11 comments
    6 saves
    2 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    5 Sources

    @voooooogellike imagine an alternate world. in this world, this reward-seeking behavior does generalize - when you deploy a reward hacking model, it tries to seize the most reward-like thing available and hillclimbs it, making a task up if nothing is available in this world everyone would just nod their head, right? nobody would be confused. "RL misalignment generalizes just like RL capabilities -- you have to be careful with reward hacking. why would you expect the coding gains to transfer and not the other behavioral changes? sounds like wishful thinking." why is that not our world?
    @xlr8harderyeah may be the weird generalizations paper is the wrong example, I agree most of the generalizations actually make quite a bit of sense, once you think about the axis on which generalizing occurs. Generalizing about authors behind text is, from one perspective, probably the most in-domain thing these models do. I still think there might be something to this, though. What do you think about the failures to generalize cross-lingually with multilingual models.

    5 Sources

    @voooooogellike imagine an alternate world. in this world, this reward-seeking behavior does generalize - when you deploy a reward hacking model, it tries to seize the most reward-like thing available and hillclimbs it, making a task up if nothing is available in this world everyone would just nod their head, right? nobody would be confused. "RL misalignment generalizes just like RL capabilities -- you have to be careful with reward hacking. why would you expect the coding gains to transfer and not the other behavioral changes? sounds like wishful thinking." why is that not our world?
    @xlr8harderyeah may be the weird generalizations paper is the wrong example, I agree most of the generalizations actually make quite a bit of sense, once you think about the axis on which generalizing occurs. Generalizing about authors behind text is, from one perspective, probably the most in-domain thing these models do. I still think there might be something to this, though. What do you think about the failures to generalize cross-lingually with multilingual models.