Models Trained on Self-Explanations Show Generalization
Adam Karvonen posted results from training models on self-generated explanations of in-the-wild behaviors.
TLDR
Adam Karvonen described training models on thousands of explanations of their own behaviors such as ignoring requests or coding errors. The single general dataset produced generalization to held-out evaluations. John Schulman called the work timely given declining CoT monitorability and noted that a metric for explanation quality enables hill-climbing via counterfactual simulatability. Laura Ruis linked related work from her group. Owain Evans amplified the thread. The posts present the training results and positive reactions without further external confirmation.
Combined views
113.1K
6 Sources, first seen 26d ago