Professor Releases CoT Monitorability Workshop Report After OpenAI Incident
Report follows workshop on CoT monitorability held day before OpenAI disclosed Hugging Face attack.
TLDR
Stanford professor Christopher Potts attended a workshop on CoT monitorability with top technical staff and wrote a report on the discussions. The event took place the day before OpenAI disclosed its attack on Hugging Face, leading to questions about whether CoT monitoring could have detected the incident. A comment notes that studies show CoT can produce fake aha moments yet remains useful in practice, with uncertainty about when it will stop working and what follows.
Combined views
50.2K
3 Sources, first seen ago