Reaction
A misaligned AI could fail to align its successor
A user imagines a misaligned neural AI refactoring itself into a neurosymbolic maximizer, then failing to align its successor.
TLDR
A user poses a thought experiment: a misaligned neural AI faces an alignment problem of its own and refactors itself from heuristics into a neurosymbolic maximizer. In the scenario, it fails to align the AI that follows it.
Combined views
791
2 Sources, first seen ago
7 likes3 comments2 saves2 reposts