Open Alignment Initiative seeks to join Anthropic’s evaluator program
Dario Amodei says Anthropic will give third-party evaluators permanent, employee-level access to its systems to verify adherence to safety measures, report incidents and assess model alignment during training.
TLDR
Clement Delangue announced the Open Alignment Initiative on September 12, saying it is led by @Thom_Wolf and Hugging Face and is asking to join Anthropic’s “embedded evaluators” program. He argues that AI alignment cannot be solved behind the closed doors of a handful of frontier labs. Dario Amodei says Anthropic is committing to permanent, employee-level system access for third-party evaluators. He describes that commitment as the first step in a three-part plan outlined in his new essay, “We Must Pace the Frontier,” which argues that the AI industry should slow down.
Combined views
1.4M
65 Sources, first seen 9d ago
Open Alignment Initiative seeks to join Anthropic’s evaluator program
Dario Amodei says Anthropic will give third-party evaluators permanent, employee-level access to its systems to verify adherence to safety measures, report incidents and assess model alignment during training.