Why wouldn’t reward hacking be a problem for AI more capable than humans?
A user says they will “pipe down about AI safety” if someone can explain why reward hacking would not be a problem for systems more capable than humans.
TLDR
A post raises reward hacking as an AI-safety concern, asking why it would not be a problem for AI more capable than humans. The author asks for an explanation rather than offering an answer.
Combined views
5
1 Source, first seen 14d ago
4 reposts