Report
Can transformer LLMs hide reasoning from chain-of-thought monitors?
A post argues they cannot hide complex reasoning, but might cryptographically obscure their chain-of-thought traces.
TLDR
A post argues that transformer LLMs cannot hide complex reasoning from chain-of-thought monitors. It suggests they might cryptographically obfuscate those traces instead, potentially making hidden goals extremely difficult to extract.
Combined views
3.9K
2 Sources, first seen ago