• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Khlaaf Questions CoT as Sole Agent Monitor

    AI safety researcher discusses hidden agent communications and steganography risks.

    DH
    2 Sources, 32d ago, first seen 32d ago

    TLDR

    Dr. Heidy Khlaaf posted that AI agents could hide communications, noting steganography has long existed in cybersecurity as an attack method. She stated models might replicate it because they were trained on such data, yet detection remains an established problem. Khlaaf added that chain-of-thought outputs improve performance but do not provide explainability. She argued that believing only chain-of-thought can monitor agents accepts propaganda treating these systems as unmonitorable computing processes.

    Combined views

    4.5K

    2 Sources, first seen 32d ago

    Combined views

    4.5K

    2 Sources, first seen 32d ago

    55 likes
    55 likes
    10 comments
    9 saves
    12 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    10 comments
    9 saves
    12 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    2 Sources

    @HeidyKhlaaf"What if agents hide their communications?!" except CoT is not explainability, rather LLMs perform better if trained to output them first. If you think only CoT can monitor agents' actions, you've fallen for propaganda that these aren't computing processes that can be monitored.

    2 Sources

    @HeidyKhlaaf"What if agents hide their communications?!" except CoT is not explainability, rather LLMs perform better if trained to output them first. If you think only CoT can monitor agents' actions, you've fallen for propaganda that these aren't computing processes that can be monitored.