Evan Hubinger Confirms Prompted and SDF'd Settings Used
Reply addresses technical questions on prior AI model experiment details.
TLDR
Evan Hubinger leads Anthropic's Alignment Stress-Testing team and researches AI alignment including inner alignment and model organisms of misalignment. In a reply to @voooooogel @BronsonSchoen and @nostalgebraist he stated that the relevant prior work had both a prompted and SDF'd setting. He added that it was a Sonnet-class model. The clarification responded to a technical question about details from earlier AI model experiments. The public posts show only this direct statement from the first-party account on the settings employed.
Combined views
73
1 Source, first seen 30d ago