Better oversight methods versus more evaluator organizations for frontier AI
The post invites collaborators to develop oversight that can scale, including ways to verify AI agents’ behavior and study collusion and covert communication in systems with multiple agents.
TLDR
A call for collaborators argues that adding evaluator organizations is not enough for effective frontier AI oversight. Topics of interest include formal and runtime verification, statistical guarantees for oversight, and supervision of very large groups of AI agents. The post also highlights collusion, covert communication and emerging cooperative norms among agents, alongside alternatives to monitoring an AI’s chain of thought and methods for AI control. It links to a form for people interested in working together.
Combined views
6K
4 Sources, first seen 18d ago