@scaling01 Advocates Defense in Depth for CoT Monitoring
Pseudonymous AI commentator @scaling01 shares views on monitoring techniques for advanced models.
TLDR
@scaling01, who runs the LisanBench LLM reasoning benchmark and posts technical analysis of model scaling, stated that defense in depth requires developing and stress-testing white-box techniques. The post added that allowing substantial degradation of CoT monitorability would be irresponsible until a roughly-as-good alternative is confirmed. The remark appears in a discussion among AI capability and safety observers on the platform.
Combined views
319.8K
37 Sources, first seen 29d ago