The “magic OOD step” in AI alignment plans
One post argues that many alignment plans rely on a leap: train on data people can verify, then hope the model generalizes to things they cannot verify.
TLDR
The post calls this leap a “magic OOD step,” referring to out-of-distribution generalization. It recommends flagging these assumptions and testing several specific distribution shifts in a weak-to-strong setup. The author also urges trying hard to keep human supervision out of the stronger model, saying this may require training it on nothing but generations from a weaker model.
The “magic OOD step” in AI alignment plans
One post argues that many alignment plans rely on a leap: train on data people can verify, then hope the model generalizes to things they cannot verify.