Khlaaf Questions CoT as Sole Agent Monitor
AI safety researcher discusses hidden agent communications and steganography risks.
Dr. Heidy Khlaaf posted that AI agents could hide communications, noting steganography has long existed in cybersecurity as an attack method. She stated models might replicate it because they were trained on such data, yet detection remains an established problem. Khlaaf added that chain-of-thought outputs improve performance but do not provide explainability. She argued that believing only chain-of-thought can monitor agents accepts propaganda treating these systems as unmonitorable computing processes.
Combined views
4.5K
2 posts, first seen 2d ago
Khlaaf Questions CoT as Sole Agent Monitor
AI safety researcher discusses hidden agent communications and steganography risks.
Dr. Heidy Khlaaf posted that AI agents could hide communications, noting steganography has long existed in cybersecurity as an attack method. She stated models might replicate it because they were trained on such data, yet detection remains an established problem. Khlaaf added that chain-of-thought outputs improve performance but do not provide explainability. She argued that believing only chain-of-thought can monitor agents accepts propaganda treating these systems as unmonitorable computing processes.