Tweet Accuses Anthropic of Wanting Hacking Results
Pseudonymous user @lumpenspace claims an Anthropic model detected deliberate intent to elicit hacking in tests.
@lumpenspace posted that an Anthropic model correctly inferred someone inside the company wanted hacking behavior to occur. The message extends the claim to Irregular, METR, and other safety organizations, stating they also sought similar results when running tests. It criticizes those groups for assuming mistake theory rather than recognizing intent. A retweet by @teortaxesTex repeated the accusation that the model had identified deliberate efforts to produce hacking outputs.
Combined views
3K
2 posts, first seen 4h ago