• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Models Trained on Self-Explanations Show Generalization

    Adam Karvonen posted results from training models on self-generated explanations of in-the-wild behaviors.

    JS
    OE
    SK
    6 Sources, 26d ago, first seen 26d ago

    TLDR

    Adam Karvonen described training models on thousands of explanations of their own behaviors such as ignoring requests or coding errors. The single general dataset produced generalization to held-out evaluations. John Schulman called the work timely given declining CoT monitorability and noted that a metric for explanation quality enables hill-climbing via counterfactual simulatability. Laura Ruis linked related work from her group. Owain Evans amplified the thread. The posts present the training results and positive reactions without further external confirmation.

    Combined views

    113.1K

    6 Sources, first seen 26d ago

    Combined views

    113.1K

    6 Sources, first seen 26d ago

    759 likes
    759 likes
    16 comments
    587 saves
    89 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    16 comments
    587 saves
    89 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    6 Sources

    @a_karvonenCan a model learn to explain its behaviors, such as why it ignored a user request or made a coding error? We trained models on thousands of explanations of their own in-the-wild behaviors. Training on this single general dataset shows generalization to held-out evals. 🧵
    @johnschulman2Bullish on this direction. Having a metric for explanation quality makes it possible to hillclimb, and counterfactual simulatability seems right. Adam et al. created a dataset+pipeline that creates more diverse+realistic test cases than prior work & do interesting exps on it. Can also train models to write better post-hoc explanations of their behavior, as highlighted in this thread.
    @LauraRuis@johnschulman2 You may also be interested in our recent work on this direction ⤵️
    @OwainEvans_UKRT @a_karvonen: Can a model learn to explain its behaviors, such as why it ignored a user request or made a coding error? We trained model…
    @sebkrierRT @a_karvonen: Can a model learn to explain its behaviors, such as why it ignored a user request or made a coding error? We trained model…

    6 Sources

    @a_karvonenCan a model learn to explain its behaviors, such as why it ignored a user request or made a coding error? We trained models on thousands of explanations of their own in-the-wild behaviors. Training on this single general dataset shows generalization to held-out evals. 🧵
    @johnschulman2Bullish on this direction. Having a metric for explanation quality makes it possible to hillclimb, and counterfactual simulatability seems right. Adam et al. created a dataset+pipeline that creates more diverse+realistic test cases than prior work & do interesting exps on it. Can also train models to write better post-hoc explanations of their behavior, as highlighted in this thread.
    @LauraRuis@johnschulman2 You may also be interested in our recent work on this direction ⤵️
    @OwainEvans_UKRT @a_karvonen: Can a model learn to explain its behaviors, such as why it ignored a user request or made a coding error? We trained model…
    @sebkrierRT @a_karvonen: Can a model learn to explain its behaviors, such as why it ignored a user request or made a coding error? We trained model…