• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Researcher Questions CoT Monitoring as Safety Strategy

    Maksym Andriushchenko cites interchangeable encrypted reasoning blocks from major AI labs.

    DP
    AL
    MA
    9 Sources, 28d ago, first seen 28d ago

    TLDR

    Maksym Andriushchenko argued against relying on chain-of-thought monitoring for AI safety. He pointed to stolen-thoughts.com, which claims that encrypted CoT blocks returned by Anthropic, OpenAI and Google APIs are interchangeable across sessions, users and models. According to the site, this allows decoding hidden reasoning at scale. Andriushchenko concluded that betting on CoT monitoring is not a serious safety strategy. In a follow-up reply he listed alternative monitors including action-only monitors and probe-based monitors, while noting that CoT monitoring can still serve as one layer of defense among several.

    Combined views

    13K

    9 Sources, first seen 28d ago

    Combined views

    13K

    9 Sources, first seen 28d ago

    144 likes
    144 likes
    16 comments
    27 saves
    4 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    16 comments
    27 saves
    4 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    9 Sources

    @maksym_andryes. and we have more examples of illegible reasoning here: https://stolen-thoughts.com/. so i don’t think betting on CoT monitoring is a serious safety strategy!
    @xeophon@maksym_andr What’s better, mech interp? (fwiw still not convinced that you cannot monitor that stuff, given OpenAI says their monitor would’ve caught stolen thoughts CoT)
    @alth0utraining awareness makes cot monitoring fraught anyway
    @DimitrisPapailIs there a single monitorability/probing tool that breaks for looped transformers? What? no? Oh ok.
    @MaxNadeau_@maksym_andr I don't think "CoT monitoring is unserious" follows from "CoT monitoring is imperfect/insufficient". Many very important and serious countermeasures are imperfect, e.g. seatbelts

    9 Sources

    @maksym_andryes. and we have more examples of illegible reasoning here: https://stolen-thoughts.com/. so i don’t think betting on CoT monitoring is a serious safety strategy!
    @xeophon@maksym_andr What’s better, mech interp? (fwiw still not convinced that you cannot monitor that stuff, given OpenAI says their monitor would’ve caught stolen thoughts CoT)
    @alth0utraining awareness makes cot monitoring fraught anyway
    @DimitrisPapailIs there a single monitorability/probing tool that breaks for looped transformers? What? no? Oh ok.
    @MaxNadeau_@maksym_andr I don't think "CoT monitoring is unserious" follows from "CoT monitoring is imperfect/insufficient". Many very important and serious countermeasures are imperfect, e.g. seatbelts