Report
Bringing AI safety evaluators into model training raises access concerns
The Information says Apollo Research wants evaluators involved during training, while safety researchers worry embedded evaluators may lack needed access.
TLDR
The Information says AI models increasingly know when they’re being tested. Apollo Research wants evaluators involved during training, rather than only checking finished models before release. The Information also describes interest in embedding outside safety researchers at AI companies, alongside concerns that embedded evaluators may not get the access they need.
Combined views
2.5K
1 Source, first seen ago
2 likes1 comments1 saves3 reposts