Work on Petri Audit Realism Shared
Target model critiques auditor actions to assess realism before selection.
TLDR
Akbir Khan retweeted a post from jplhughes announcing new work on Petri audit realism. The post states that the target model first critiques the auditor actions for realism and then picks which actions to evaluate further. Khan works on AI alignment at Anthropic with a focus on scalable oversight and AI debate. The announcement appears in a thread tagged for AI safety topics. No additional details on results or methods appear in the visible post. The retweet highlights the project within ongoing conversations about auditor evaluation techniques.
Combined views
1 Source, first seen 24d ago