Steering AI with opposing fine-tunes may often generalize better than activation steering
A post says the method uses the weight difference between two opposite fine-tunes as a steering direction.
TLDR
A post describing the authors’ work says steering along the weight difference between opposite fine-tunes is often more generalizable than activation steering. It also quotes a finding that emergent misalignment can be detected by comparing fine-tuning updates with an “evil” weight direction.
Combined views
2
1 Source, first seen ago