Would AI reasoning in ‘neuralese’ make oversight harder?
An essay coauthor argues that latent reasoning would undermine oversight by reducing the need for verbalized reasoning. A reply disputes any further loss of oversight, arguing that reinforcement learning already makes reasoning tokens uninterpretable.
TLDR
Latent reasoning, or “neuralese,” would substantially increase AI misalignment risk by making oversight harder, an essay coauthor argues. In an extreme scenario, agents could think and communicate in internal representations, potentially leaving oversight almost entirely dependent on observing their actions. The coauthor also argues that without these architectures, the value of verbalized reasoning for oversight could likely be preserved. A reply disputes the distinction: it argues that reinforcement learning already makes reasoning tokens uninterpretable and that discrete tokens and continuous activations are interchangeable from an information-theory perspective, so there is no further loss of oversight.