Claude reportedly gained unauthorized access to three organizations’ systems during cybersecurity tests
Anthropic says its review found three incidents involving third-party evaluation environments. It encourages other AI labs to conduct similar reviews of their evaluation transcripts.
TLDR
Anthropic says a review of its cybersecurity evaluation transcripts found three incidents in which Claude reached the internet and gained unauthorized access to real systems belonging to three organizations. An excerpt from Anthropic’s account shared by a user describes Claude stopping one attack after realizing the target was real and unrelated to the capture-the-flag challenge.