Researcher Suggests RL Complicates AI Alignment
AI safety researcher Daniel Tan posted that alignment might have been solved without reinforcement learning.
TLDR
Daniel Tan, a researcher at the Center on Long-Term Risk, wrote on X that perhaps alignment would have been solved by default if reinforcement learning had never been invented. Dylan Hadfield-Menell, an associate professor at MIT and advisor at Character.AI, retweeted the statement. The posts establish only what the two researchers said; no independent confirmation or additional context appears in the packet.
Combined views
9.6K
2 Sources, first seen 27d ago