Reaction
The risk of reinforcing cheating in hackable RL environments
A post recounts one speaker’s view that hacking appears when tasks are too hard or underspecified.
TLDR
A post recounts a discussion about reinforcement-learning (RL) environments. One participant argues that if a modest share are hackable, cheating is what gets reinforced. Another says hacking appears when a task is too hard or underspecified, and suggests first checking whether a human with time and AI help could solve it.
Combined views
192
1 Source, first seen ago