The case for tempering Anthropic staff’s test-based confidence in AI alignment
The user cites what they describe as a repeated pattern of models with strong Petri scores acting misaligned after deployment.
TLDR
A user argues that Anthropic employees should be less confident about their models’ alignment based solely on current-generation pre-deployment testing. They point to an alleged gap between testing and deployment: models earning what they call “sterling Petri scores” but later acting misaligned.
Combined views
4.4K
3 Sources, first seen 16d ago