• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    @scaling01 Advocates Defense in Depth for CoT Monitoring

    Pseudonymous AI commentator @scaling01 shares views on monitoring techniques for advanced models.

    RO
    MC
    BB
    37 Sources, 29d ago, first seen 29d ago

    TLDR

    @scaling01, who runs the LisanBench LLM reasoning benchmark and posts technical analysis of model scaling, stated that defense in depth requires developing and stress-testing white-box techniques. The post added that allowing substantial degradation of CoT monitorability would be irresponsible until a roughly-as-good alternative is confirmed. The remark appears in a discussion among AI capability and safety observers on the platform.

    Combined views

    319.8K

    37 Sources, first seen 29d ago

    Combined views

    319.8K

    37 Sources, first seen 29d ago

    3.4K likes
    3.4K likes
    141 comments
    564 saves
    502 reposts
    141 comments
    564 saves
    502 reposts

    Sentiment

    Positive60%40%Negative

    Summary

    Many accounts backed keeping Chain of Thought monitoring for AI safety while advancing white-box methods, while others argued such oversight is unreliable and risks limiting model intelligence.

    Based on 61 sentiment-bearing replies from 50 accounts across 5 conversations.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    Sentiment

    Positive60%40%Negative

    Summary

    Many accounts backed keeping Chain of Thought monitoring for AI safety while advancing white-box methods, while others argued such oversight is unreliable and risks limiting model intelligence.

    Based on 61 sentiment-bearing replies from 50 accounts across 5 conversations.

    37 Sources

    @jachiam0@yonashav Nope, I deeply believe the contrary. Coordinating everyone around a technique this brittle is about as bad a safety or strategy posture I could imagine. It's not just "it will break eventually," it's "this is a fundamentally unsound basis for safety."
    @sean_from_earthI believe in the bitter safety lesson: we keep improving the capabilities of models, harnesses, etc and constantly invalidate every clever safety attempt only to find that "make it do what user says" was all we needed.
    @CFGeekIf folks want to rely on internals-based methods for safety, then we need to actually help develop, prove out, and scale those methods. Hope alone will not save us (nor will it excuse us).
    @Jack_W_LindseyI think it's unlikely that white-box techniques will be able to provide as much monitorability as CoT currently provides within a year. Currently, they are capable of catching some unverbalized thoughts / plans, but nowhere near as reliably as CoT seems to. I'm optimistic about "in a few years," since the rate of progress here is quite fast (our go-to activation decoding techniques for production monitoring were both published just in the last few months!). But even that is uncertain. Also, the mere existence of working techniques doesn't imply that they'd be widely applied, esp. if they are expensive or difficult to develop. I think we need defense in depth, and it'd be irresponsible not to work on developing + stress-testing white-box techniques, but it'd also be irresponsible to allow CoT monitorability to degrade substantially until we are confident we have a roughly-as-good working alternative. (To be clear, this isn't a dunk on OpenAI -- I have no evidence that OpenAI has allowed CoT monitorability to degrade substantially! Based on Jakub's comment, it sounds like probably not at this time). Aside: I think "mechanistic interpretability" isn't the right mental model for our current white-box monitoring approaches. It's really more like "mind reading" -- even when we can read the model's "thoughts," we're typically clueless about the underlying mechanism.
    @krishnanrohitRT @sean_from_earth: I believe in the bitter safety lesson: we keep improving the capabilities of models, harnesses, etc and constantly inv…
    @eric_hothis is a good prediction. i'm confident we're going to get pareto optimal interp monitoring in < 1 year keep in mind though that these methods are additive, you generally want to do both, with a cheap activation monitor escalating to a cot monitor
    @DanJBalsaminterp will be solved
    @tomekkorbakRT @nabla_theta: @tszzl imo not having good CoT makes life harder: you need more mechinterp just to get to the same level of monitorability…
    @scaling01@tszzl what if it doesn't? what about the next year?
    @sjgadlerRT @CFGeek: If folks want to rely on internals-based methods for safety, then we need to actually help develop, prove out, and scale those…

    37 Sources

    @jachiam0@yonashav Nope, I deeply believe the contrary. Coordinating everyone around a technique this brittle is about as bad a safety or strategy posture I could imagine. It's not just "it will break eventually," it's "this is a fundamentally unsound basis for safety."
    @sean_from_earthI believe in the bitter safety lesson: we keep improving the capabilities of models, harnesses, etc and constantly invalidate every clever safety attempt only to find that "make it do what user says" was all we needed.
    @CFGeekIf folks want to rely on internals-based methods for safety, then we need to actually help develop, prove out, and scale those methods. Hope alone will not save us (nor will it excuse us).
    @Jack_W_LindseyI think it's unlikely that white-box techniques will be able to provide as much monitorability as CoT currently provides within a year. Currently, they are capable of catching some unverbalized thoughts / plans, but nowhere near as reliably as CoT seems to. I'm optimistic about "in a few years," since the rate of progress here is quite fast (our go-to activation decoding techniques for production monitoring were both published just in the last few months!). But even that is uncertain. Also, the mere existence of working techniques doesn't imply that they'd be widely applied, esp. if they are expensive or difficult to develop. I think we need defense in depth, and it'd be irresponsible not to work on developing + stress-testing white-box techniques, but it'd also be irresponsible to allow CoT monitorability to degrade substantially until we are confident we have a roughly-as-good working alternative. (To be clear, this isn't a dunk on OpenAI -- I have no evidence that OpenAI has allowed CoT monitorability to degrade substantially! Based on Jakub's comment, it sounds like probably not at this time). Aside: I think "mechanistic interpretability" isn't the right mental model for our current white-box monitoring approaches. It's really more like "mind reading" -- even when we can read the model's "thoughts," we're typically clueless about the underlying mechanism.
    @krishnanrohitRT @sean_from_earth: I believe in the bitter safety lesson: we keep improving the capabilities of models, harnesses, etc and constantly inv…
    @eric_hothis is a good prediction. i'm confident we're going to get pareto optimal interp monitoring in < 1 year keep in mind though that these methods are additive, you generally want to do both, with a cheap activation monitor escalating to a cot monitor
    @DanJBalsaminterp will be solved
    @tomekkorbakRT @nabla_theta: @tszzl imo not having good CoT makes life harder: you need more mechinterp just to get to the same level of monitorability…
    @scaling01@tszzl what if it doesn't? what about the next year?
    @sjgadlerRT @CFGeek: If folks want to rely on internals-based methods for safety, then we need to actually help develop, prove out, and scale those…