The case for external checks and access control in AI alignment
A commenter worries that finding isolated features in AI models could look promising while creating false trust. They argue instead for clear checks, traditional software and access controls.
TLDR
In a series of replies, one commenter argues that AI alignment needs a mostly external “immune system” of clear symbolic checks, traditional software and access controls. They liken the approach to law, which they describe as access control with humans involved where it matters. Even if “RL superalignment” worked, they say, disagreements would remain and people would still need to build a shared framework.
Combined views
3.1K
9 Sources, first seen 10h ago