LessWrong Post Flags Gaps in OpenAI Monitorability Evals
Sarah Wiegreffe shares a LessWrong post examining limits in chain-of-thought monitoring tests.
Sarah Wiegreffe posted a LessWrong analysis titled Evaluating Chain-of-Thought Monitorability is Still an Open Problem. The piece comments on OpenAI's monitorability evaluations and argues that current methods for testing chain-of-thought oversight leave key questions unresolved. Christopher Potts called the work timely and valuable. The post draws on feedback from Iván Arcuschin Moreno and focuses on research problems that remain open rather than settled findings.
Combined views
7.6K
2 posts, first seen 2d ago