• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Evan Hubinger Confirms Prompted and SDF'd Settings Used

    Reply addresses technical questions on prior AI model experiment details.

    EH
    1 Source, 30d ago, first seen 30d ago

    TLDR

    Evan Hubinger leads Anthropic's Alignment Stress-Testing team and researches AI alignment including inner alignment and model organisms of misalignment. In a reply to @voooooogel @BronsonSchoen and @nostalgebraist he stated that the relevant prior work had both a prompted and SDF'd setting. He added that it was a Sonnet-class model. The clarification responded to a technical question about details from earlier AI model experiments. The public posts show only this direct statement from the first-party account on the settings employed.

    Combined views

    73

    1 Source, first seen 30d ago

    Combined views

    73

    1 Source, first seen 30d ago

    4 likes
    4 likes
    1 comments
    1 saves

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    1 comments
    1 saves

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    1 Source

    @EvanHub@voooooogel @BronsonSchoen @nostalgebraist We had both a prompted and SDF'd setting. And it was a Sonnet-class model.

    1 Source

    @EvanHub@voooooogel @BronsonSchoen @nostalgebraist We had both a prompted and SDF'd setting. And it was a Sonnet-class model.