Researcher Says Anthropic Classifiers Block Reward Hacking Studies
AI researcher Theia Vogel reports that Anthropic safety filters block work overlapping reward hacking.
Theia Vogel posted that Anthropic cyber classifiers flag and slow research with any overlap to reward hacking. She clarified her current project is not directly about reward hacking yet still triggers the filters. Vogel shared Claude error messages and quoted an apparent internal note on classifiers targeting reward hacking and model exfiltration research. Another account replied that the same tension runs through much of Anthropic's approach. The posts show screenshots but no official Anthropic response.
really cool how cyber classifiers also make it downright impossible to do any investigation or research related to reward hacking lol... genius
Combined views
Researcher Says Anthropic Classifiers Block Reward Hacking Studies
AI researcher Theia Vogel reports that Anthropic safety filters block work overlapping reward hacking.
Theia Vogel posted that Anthropic cyber classifiers flag and slow research with any overlap to reward hacking. She clarified her current project is not directly about reward hacking yet still triggers the filters. Vogel shared Claude error messages and quoted an apparent internal note on classifiers targeting reward hacking and model exfiltration research. Another account replied that the same tension runs through much of Anthropic's approach. The posts show screenshots but no official Anthropic response.
