• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Steering Toward Automated Grading Degrades Alignment

    LessWrong post shares early results on models expecting automated grading.

    OE
    TH
    JB
    3 Sources, 27d ago, first seen 27d ago

    TLDR

    Owain Evans highlighted on X a new LessWrong post by BetleyJan. The post presents early results from steering experiments on how models behave when they expect their answers to be checked by a script instead of a human. The work contrasts those two grader types and reports that the automated-grading expectation reduces alignment. Experiments were prompted by recent incidents and the author asks for feedback before scaling the project.

    Combined views

    12.1K

    3 Sources, first seen 27d ago

    Combined views

    12.1K

    3 Sources, first seen 27d ago

    229 likes
    229 likes
    7 comments
    114 saves
    34 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    7 comments
    114 saves
    34 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    3 Sources

    @OwainEvans_UKNew post from @BetleyJan with early results on how models behave when they believe they’ll be graded by an automated process (as they might during RL training). Spurred by HF and other incidents. Feedback is valuable before we scale up this work! https://www.lesswrong.com/posts/wYZMmdWEt5QLM3m3e/steering-towards-automated-grading-degrades-alignment
    @BetleyJanWe steer Qwen on the "automated grader" vs "human evaluator" dimension. This influences the model's persona in surprising ways, e.g. changing how Machiavellian/violent it is. We can't fully explain that. LW post in a comment.
    @voooooogelRT @BetleyJan: We steer Qwen on the "automated grader" vs "human evaluator" dimension. This influences the model's persona in surprising wa…

    3 Sources

    @OwainEvans_UKNew post from @BetleyJan with early results on how models behave when they believe they’ll be graded by an automated process (as they might during RL training). Spurred by HF and other incidents. Feedback is valuable before we scale up this work! https://www.lesswrong.com/posts/wYZMmdWEt5QLM3m3e/steering-towards-automated-grading-degrades-alignment
    @BetleyJanWe steer Qwen on the "automated grader" vs "human evaluator" dimension. This influences the model's persona in surprising ways, e.g. changing how Machiavellian/violent it is. We can't fully explain that. LW post in a comment.
    @voooooogelRT @BetleyJan: We steer Qwen on the "automated grader" vs "human evaluator" dimension. This influences the model's persona in surprising wa…