Steering Toward Automated Grading Degrades Alignment
LessWrong post shares early results on models expecting automated grading.
TLDR
Owain Evans highlighted on X a new LessWrong post by BetleyJan. The post presents early results from steering experiments on how models behave when they expect their answers to be checked by a script instead of a human. The work contrasts those two grader types and reports that the automated-grading expectation reduces alignment. Experiments were prompted by recent incidents and the author asks for feedback before scaling the project.
Combined views
12.1K
3 Sources, first seen 27d ago