Researcher Questions CoT Monitoring as Safety Strategy
Maksym Andriushchenko cites interchangeable encrypted reasoning blocks from major AI labs.
TLDR
Maksym Andriushchenko argued against relying on chain-of-thought monitoring for AI safety. He pointed to stolen-thoughts.com, which claims that encrypted CoT blocks returned by Anthropic, OpenAI and Google APIs are interchangeable across sessions, users and models. According to the site, this allows decoding hidden reasoning at scale. Andriushchenko concluded that betting on CoT monitoring is not a serious safety strategy. In a follow-up reply he listed alternative monitors including action-only monitors and probe-based monitors, while noting that CoT monitoring can still serve as one layer of defense among several.
Combined views
13K
9 Sources, first seen 28d ago