Trask Says Alignment Problems Trace to Training Data
DeepMind researcher argues data problems cause alignment issues and fixes.
TLDR
Andrew Trask posted that the main sources of AI alignment failures are unconstrained training data sets carrying strong dangerous signals, especially web scrapes and RL reward mechanisms. He added that the largest alignment advances consist of data interventions such as filtering, RLHF, and reward shaping that remove or override those signals. Trask stated that researchers need direct access to training and RL data records to make progress, yet most lack this access and therefore work without visibility into the source of the behaviors they try to correct. He proposed structured transparency tools to expand the number of people who can study the data safely.
Combined views
17K
7 Sources, first seen 26d ago