Anthropic safety team member announces departure, urges more AI transparency
The former team member says they plan to join METR to evaluate AI risks independently, arguing that AI companies are underinvesting in safety.
TLDR
On September 11, a former member of Anthropic's safety team said they had left two weeks earlier and planned to join METR for independent AI risk evaluations. They warned that a company could lose control of its systems without the public knowing. Their proposed guardrails include disclosing progress toward recursive self-improvement—AI systems improving themselves—reporting safety incidents and near-misses, and meeting minimum safety standards with independent guarantees that those standards are met.
Combined views
2.1M
6 Sources, first seen 19d ago