Research Makes Petri Audits Harder for Models to Detect
@_axelahlqvist posted about a new paper on alignment audits.
TLDR
Ethan Perez retweeted a post from @_axelahlqvist. The post states that alignment audits only work if the model cannot tell it is being audited. It announces a new paper that makes Petri audits far more resistant to detection. The message cuts off after the word more. Perez leads the alignment and adversarial robustness team at Anthropic. The packet contains no additional details from the paper itself or from other sources.
Combined views
15
1 Source, first seen 24d ago