• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Stanford Researcher Warns Deceptive Chain-Of-Thought Requires Internal Monitoring

    Stanford professor argues chain-of-thought outputs cannot be trusted as faithful reflections of model computation.

    Christopher PottsCP
    Pasquale MinerviniPM
    2 Sources, 70d ago, first seen 70d ago

    TLDR

    Christopher Potts, Stanford Professor and Chair of Linguistics, addressed the threat of deceptive chain-of-thought in Section 3 of recent work. He stressed that CoT cannot be assumed to faithfully reflect a model's actual computation and called for monitoring internal states or detecting subliminal-learning-style effects. A linked discussion highlighted a blog post examining CoT monitorability, with Pasquale Minervini noting its relevance for NLP and AI safety researchers tracking reasoning transparency.

    Combined views

    2.6K

    2 Sources, first seen 70d ago

    Combined views

    2.6K

    2 Sources, first seen 70d ago

    15 likes
    15 likes
    1 comments
    4 saves
    1 reposts
    1 comments
    4 saves
    1 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    2 Sources

    Christopher Potts@ChrisGPottsSection 3: "On the threat of deceptive CoT". This concerns the importance of not trusting CoT to be a faithful reflection of a model's computation. I think this too argues for monitoring internal states, or at least looking for subliminal-learning-style effects in CoT.70d
    Pasquale Minervini@PMinervinisuper nice blog post on CoT monitorability69d

    2 Sources

    Christopher Potts@ChrisGPottsSection 3: "On the threat of deceptive CoT". This concerns the importance of not trusting CoT to be a faithful reflection of a model's computation. I think this too argues for monitoring internal states, or at least looking for subliminal-learning-style effects in CoT.70d
    Pasquale Minervini@PMinervinisuper nice blog post on CoT monitorability69d
    Christopher Potts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Related

    Naomi Saphra Writes on AI and the Uncanny

    Researcher Naomi Saphra shares an essay on dissociation triggered by realistic AI outputs.

    Hugging Face Creates Opportunity To Learn From Event And Prevent Future Risks
    Denial Of Knowledge Does Not Yield Safer AI Models, Section Argues