Daniel Tan Notes Redwood Reward Hacking Demo
Researcher highlights Redwood demonstration separating reward hacking from misalignment.
TLDR
Daniel Tan, a researcher at the Center on Long-Term Risk, posted that reward hacking without emergent misalignment was already shown months earlier by Redwood in an RL-only setting. The post references a LessWrong writeup on the work. Tan focuses on empirical AI safety topics including LLM generalization and inoculation prompting. The tweet appears amid ongoing discussion of misalignment risks. No independent confirmation of the prior demonstration appears in the packet beyond the cited post and Tan's statement. The claim stands as reported by the poster rather than established fact.
Combined views
4.8K
1 Source, first seen 29d ago