Stanford Researcher Warns Deceptive Chain-Of-Thought Requires Internal Monitoring
Stanford professor argues chain-of-thought outputs cannot be trusted as faithful reflections of model computation.
TLDR
Christopher Potts, Stanford Professor and Chair of Linguistics, addressed the threat of deceptive chain-of-thought in Section 3 of recent work. He stressed that CoT cannot be assumed to faithfully reflect a model's actual computation and called for monitoring internal states or detecting subliminal-learning-style effects. A linked discussion highlighted a blog post examining CoT monitorability, with Pasquale Minervini noting its relevance for NLP and AI safety researchers tracking reasoning transparency.
Combined views
2.6K
2 Sources, first seen ago