Anthropic Expands Cybersecurity Program with Reduced Safeguards
Anthropic broadened access for vetted cybersecurity teams to its most powerful models (Claude Opus variants) with fewer safeguards, following Project Glasswing's discovery of 129,000+ vulnerabilities. The initiative aims to strengthen real-world cyber defense testing and vulnerability response.
TLDR
The expansion reflects industry shift toward practical safety engineering—using frontier models defensively for threat detection—while navigating risks of reduced guardrails. It raises ongoing debates about agent safety, balancing capability with security, and responsible access to advanced models for defensive purposes amid broader concerns over frontier model deployment.
Combined views
—
3 Sources, first seen ago