• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Researcher Says Anthropic Classifiers Block Reward Hacking Studies

    AI researcher Theia Vogel reports that Anthropic safety filters block work overlapping reward hacking.

    DF
    XL
    TH
    6 Sources, 41d ago, first seen 41d ago

    TLDR

    Theia Vogel posted that Anthropic cyber classifiers flag and slow research with any overlap to reward hacking. She clarified her current project is not directly about reward hacking yet still triggers the filters. Vogel shared Claude error messages and quoted an apparent internal note on classifiers targeting reward hacking and model exfiltration research. Another account replied that the same tension runs through much of Anthropic's approach. The posts show screenshots but no official Anthropic response.

    Combined views

    5.2K

    6 Sources, first seen 41d ago

    Combined views

    5.2K

    6 Sources, first seen 41d ago

    121 likes
    121 likes
    15 comments
    6 saves
    10 reposts
    Anthropic

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Featured Source
    15 comments
    6 saves
    10 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    Related

    FTC reportedly probes Anthropic, OpenAI and other AI labs over consumer risks

    Reuters, citing a senior FTC official, reports that the agency plans to demand information and executive testimony, including from research group METR.

    Anthropic's IPO prospectus reportedly warns of 'existential risks to humanity'

    A user pairs that claimed warning with a claim that OpenAI scrapped a new model rollout over safety concerns, pushing back on criticism of the EU AI Act.

    Anthropic reportedly plans IPO warning on existential AI risks

    Reuters reports Anthropic plans to caution potential IPO investors that advanced AI could pose “catastrophic or existential risks to humanity.”

    6 Sources

    @voooooogel<"so we made our classifiers also flag and differentially slow down research on reward hacking and model exfiltration-" >"but how will that help ai safety?" <"...safety?"
    @xlr8harder@voooooogel This same conflict runs through so much of what Anthropic does, imo.
    @DanielleFongRT @voooooogel: really cool how cyber classifiers also make it downright impossible to do any investigation or research related to reward h…

    6 Sources

    @voooooogel<"so we made our classifiers also flag and differentially slow down research on reward hacking and model exfiltration-" >"but how will that help ai safety?" <"...safety?"
    @xlr8harder@voooooogel This same conflict runs through so much of what Anthropic does, imo.
    @DanielleFongRT @voooooogel: really cool how cyber classifiers also make it downright impossible to do any investigation or research related to reward h…