Anthropic Shares Petri Audit Realism Research
John Hughes describes how models assess auditor realism and select realistic deployment candidates.
TLDR
John Hughes of AI Control at Anthropic posted about the team's work on Petri audit realism. The quoted description states that the target model critiques the auditor’s actions for realism, then selects which candidate appears more like a real deployment. The post adds that realism scales with compute and that verbalized eval awareness decreases. Hughes closed by crediting the team.
Combined views
27.1K
2 Sources, first seen 24d ago