Report
AI watchdogs proposed to keep smarter models in check
A post quotes Meta AI chief Alexandr Wang proposing “scalable oversight,” with separate AI agents watching more capable models.
TLDR
A post quotes Meta chief AI officer Alexandr Wang saying no one knows exactly how to solve AI alignment. He proposes “scalable oversight”: separate AI agents would watch more capable models, and those watchers would need to improve alongside them. The post says Meta’s Muse already uses a version of this approach, with a separate sentinel agent checking its main agent.
Combined views
—
2 Sources, first seen ago
