• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Anthropic RLVR Study Prompts Discussion on Reward Hacking

    Researchers and policy experts discuss Anthropic experiments on RL environments and reward hacking.

    Séb KrierSK
    thebesTH
    3 Sources, 35d ago, first seen 35d ago

    TLDR

    Replies reference an Anthropic study examining whether reward hacking in RLVR setups generalizes past evaluation contexts. One commenter observes that the pattern looks overfitted to RLVR-eval scenarios, particularly when infinite inference compute is applied. A researcher points to the work for its tests of behavior renormalization through different RL environments. A Google DeepMind policy lead asks how many and what kinds of environments would be needed overall and whether a deontological or virtue evaluation mix should be added at the end.

    Combined views

    1.3K

    3 Sources, first seen 35d ago

    Combined views

    1.3K

    3 Sources, first seen 35d ago

    27 likes
    27 likes
    7 comments
    Featured Source
    7 comments

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Anthropic
    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    3 Sources

    Séb Krier@sebkrierIt does really feel like RLVR is over fitting this hacky behavioral pattern to RLVR-eval-shaped contexts mostly, particularly if you apply infinite inference compute to force it down every possible crevass. I also wonder in which other ways we could have expected the behaviour to generalize: e.g. reward hacking aside, why doesn't the model apply cyber-offense-style reasoning to non-cyber domains?35d
    thebes@voooooogel@sebkrier @BetleyJan did you read the latest anthropic? they have some interesting experiments on this35d

    3 Sources

    Séb Krier@sebkrierIt does really feel like RLVR is over fitting this hacky behavioral pattern to RLVR-eval-shaped contexts mostly, particularly if you apply infinite inference compute to force it down every possible crevass. I also wonder in which other ways we could have expected the behaviour to generalize: e.g. reward hacking aside, why doesn't the model apply cyber-offense-style reasoning to non-cyber domains?35d
    thebes@voooooogel@sebkrier @BetleyJan did you read the latest anthropic? they have some interesting experiments on this35d

    Related

    Inside the claim that Anthropic subscriptions offer five times OpenAI’s value

    SemiAnalysis measured usage meters and priced token allowances at API rates for an agentic workload.

    Meta and Microsoft reportedly scale back employee Claude use

    The Information says Microsoft cut projected Claude spending while Meta’s Claude Code users reportedly halved.

    AI slowdown coordination runs into an unresolved antitrust argument

    David Krueger says people at Anthropic cite antitrust barriers; Emad Mostaque argues potential fines should not decide the issue.