• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Paper explores probe-based training for AI alignment

    The researchers report robustness and out-of-distribution harmfulness results while keeping the model monitorable.

    X(
    MA
    JA
    10 Sources, 10h ago, first seen 10h ago

    TLDR

    Researchers describe training against internal probes to improve a model’s alignment while preserving white-box monitoring. In one harmfulness test, they say probes were trained on BeaverTails, the model was rolled out on JBB, and evaluation on ClearHarm showed out-of-distribution generalization. The team says its model remained robust and monitorable, though before-and-after monitorability results across all data splits were still to be added.

    Combined views

    13.1K

    10 Sources, first seen 10h ago

    Combined views

    13.1K

    10 Sources, first seen 10h ago

    193 likes
    193 likes
    6 comments
    101 saves
    57 reposts

    Researchers behind a new AI safety paper say training against internal model signals can improve alignment without making the model harder to inspect. In the paper announcement, Lena Libon framed the problem as one in which frontier models can appear aligned during training but later show unwanted behavior in deployment.

    The team describes the method as training against probes, using a model’s internal activity as a training signal. Ben Rank said the resulting model was more aligned and robust while remaining monitorable. He noted that some researchers call this a “forbidden technique” because of concern that it could reduce monitorability.

    A harmfulness test across datasets

    Alexander Panfilov outlined one harmfulness experiment: the researchers trained probes on BeaverTails, rolled the model out on JBB and evaluated it on ClearHarm. He said the behavior generalized to ClearHarm even though that dataset was out of distribution relative to the training and rollout data.

    The researchers have not yet supplied every monitorability comparison. Panfilov said before-and-after results for all data splits would be added, leaving an important part of the paper’s central claim pending in the material shared so far.

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Featured Source
    6 comments
    101 saves
    57 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    #17

    Today's Rank

    #17

    10 Sources

    @lenalibonFrontier models can look aligned during training while later showing unwanted behaviour in deployment. Can we use internal signals during training to align models better without making white-box monitoring harder? Our new paper suggests yes 🧵10h
    @kotekjedi_mlCheck out the extension of our MechInterp workshop submission on how alignment through training against probes could work! I think work in this direction would be very important if we are to stick with latent reasoning models10h
    @full__rankModel internals can be used for training aligned models. Some people call this the "forbidden technique", because it could decrease monitorability. In this work, we show that this is not necessarily the case. You get quite an aligned and robust model, and it is still monitorable10h
    @maksym_andrCheck out the detailed thread by @lenalibon about our new paper on using model internals for improving alignment (aka "the most forbidden technique")!9h
    @juliusadmlImportant lesson that every era of interpretability research often needs to relearn: you only get reliable understanding/control if you optimize for it. Classifier guidance works even better if you insert it inside the model :) It gives you control and understanding simultaneously.7h
    @xuanalogueRT @lenalibon: Frontier models can look aligned during training while later showing unwanted behaviour in deployment. Can we use internal…43m

    10 Sources

    @lenalibonFrontier models can look aligned during training while later showing unwanted behaviour in deployment. Can we use internal signals during training to align models better without making white-box monitoring harder? Our new paper suggests yes 🧵10h
    @kotekjedi_mlCheck out the extension of our MechInterp workshop submission on how alignment through training against probes could work! I think work in this direction would be very important if we are to stick with latent reasoning models10h
    @full__rankModel internals can be used for training aligned models. Some people call this the "forbidden technique", because it could decrease monitorability. In this work, we show that this is not necessarily the case. You get quite an aligned and robust model, and it is still monitorable10h
    @maksym_andrCheck out the detailed thread by @lenalibon about our new paper on using model internals for improving alignment (aka "the most forbidden technique")!9h
    @juliusadmlImportant lesson that every era of interpretability research often needs to relearn: you only get reliable understanding/control if you optimize for it. Classifier guidance works even better if you insert it inside the model :) It gives you control and understanding simultaneously.7h
    @xuanalogueRT @lenalibon: Frontier models can look aligned during training while later showing unwanted behaviour in deployment. Can we use internal…43m