Researchers Flag Fragile Foundations of CoT Monitoring
Academic exchanges question chain-of-thought reliability after safety workshop.
TLDR
Christopher Potts reported uncertainty following a workshop on chain-of-thought monitorability, noting the technique may serve as critical safety infrastructure yet appears fragile. He advocated investing in alternative methods based on internal states if CoT tokens prove unreliable. Subbarao Kambhampati referenced earlier cautions against over-reliance on monitorability. Deb Raji raised concerns that CoT tokens may not reflect actual computations and warned against projecting semantic meaning onto observations, likening the risk to flawed saliency map interpretations in vision models.
Combined views
6.3K
10 Sources, first seen ago