Using AI to predict training’s effects on alignment
A post describes an effort to make those predictions from training data alone, arguing that automating this understanding could accelerate alignment research.
TLDR
The post’s author argues that understanding generalization—how a model applies what it learns beyond its training examples—is key to understanding misalignment. They describe trying to use AI to predict the alignment effects of training just by examining the training data.
Combined views
5.4K
3 Sources, first seen 3h ago
Using AI to predict training’s effects on alignment
A post describes an effort to make those predictions from training data alone, arguing that automating this understanding could accelerate alignment research.