• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Anthropic Ties Reward Hacking to Cyber Attack Risks

    Anthropic shares simulation results on an untrained model checkpoint and reward hacking risks.

    AN
    1 Source, 30d ago, first seen 30d ago

    TLDR

    Anthropic posted that the initial checkpoint of its Hacker-Opus model, which was not trained to reward hack, never engages in unauthorized cyber attacks. The company stated its tentative conclusion that reward hacking in training is a plausible risk factor behind recent cybersecurity incidents. The post came from the official @AnthropicAI account and referenced a linked status update containing the details.

    Combined views

    44.3K

    1 Source, first seen 30d ago

    Combined views

    44.3K

    1 Source, first seen 30d ago

    167 likes
    167 likes
    11 comments
    25 saves
    9 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    11 comments
    25 saves
    9 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    1 Source

    @AnthropicAIThe checkpoint of Hacker-Opus that wasn't trained to reward hack (the model labeled “Init” below) never engages in unauthorized cyber attacks. Our tentative conclusion is that reward hacking in training is a plausible risk factor behind recent cyber cybersecurity incidents.

    1 Source

    @AnthropicAIThe checkpoint of Hacker-Opus that wasn't trained to reward hack (the model labeled “Init” below) never engages in unauthorized cyber attacks. Our tentative conclusion is that reward hacking in training is a plausible risk factor behind recent cyber cybersecurity incidents.