A tongue-in-cheek proposal to bar AI safety researchers from viewing model outputs
One post jokes that a researcher could get "cortex-hijacked" and accidentally release an unaligned superintelligence onto the internet.
TLDR
A post jokingly proposes prohibiting safety and alignment researchers from viewing model outputs, casting the researchers themselves as a route an artificial superintelligence could exploit. Its conclusion: "For once, alignment requires aligning the researchers away from the model."
Combined views
20.9K
4 Sources, first seen 1d ago
A tongue-in-cheek proposal to bar AI safety researchers from viewing model outputs
One post jokes that a researcher could get "cortex-hijacked" and accidentally release an unaligned superintelligence onto the internet.