@yonashav The combination of "A model has the theoretical capability of invisible steganography" and "We give incentives during training to develop it" (as we repeatedly see) is sufficient for me to believe quite firmly CoT monitoring is insufficient (even with temporary…
@yonashav CoT monitoring by itself is inherently doomed to fail (e.g., our works on computationally undetectable steganography)
Has anyone done any empirical estimations of the effect of indirect CoT pressure on CoT monitorability? E.g. by taking CoT-detected-reward-hack examples in a toy env, programmatically rewriting them to exclude hack-reasoning while preserving the misbehavior, and updating against…
see also