Reasoning traces alone versus reasoning and actions in AI monitoring
A user says their production AI monitors see both reasoning traces and actions, and work substantially better—for now—than the reasoning-only monitors they typically use in monitorability tests.
TLDR
A user says they typically measure chain-of-thought (CoT) monitorability with monitors that see only an AI's reasoning traces. In production, they use monitors that also see actions and say those work substantially better for now. The same user says, anecdotally, that their CoT monitors would have flagged the Hugging Face incident, while Anthropic's would have flagged its own incidents.
Combined views
624
2 Sources, first seen 15d ago