Researcher Says Anthropic Classifiers Block Reward Hacking Studies
AI researcher Theia Vogel reports that Anthropic safety filters block work overlapping reward hacking.
TLDR
Theia Vogel posted that Anthropic cyber classifiers flag and slow research with any overlap to reward hacking. She clarified her current project is not directly about reward hacking yet still triggers the filters. Vogel shared Claude error messages and quoted an apparent internal note on classifiers targeting reward hacking and model exfiltration research. Another account replied that the same tension runs through much of Anthropic's approach. The posts show screenshots but no official Anthropic response.
Combined views
5.2K
6 Sources, first seen 41d ago

