A new token may help confine AI misalignment to a special mode
Geodesic Research says its paper demonstrates a way to keep AI aligned outside a mode marked by a newly coined token. A coauthor calls the method preliminary.
TLDR
Geodesic Research describes midtraining models on synthetic documents explaining how AI can be misaligned in a special, token-marked mode but aligned outside it. A coauthor says they then use reinforcement learning on misaligned data with the token present, followed by evaluation without it. The coauthor says this lets them control how broadly misalignment generalizes, but calls the method preliminary.
Combined views
29.3K
9 Sources, first seen 15d ago