• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Boaz Barak Reports Astra Alignment Advances Over Sol

    Harvard professor and OpenAI alignment researcher discusses advances with Astra compared to Sol but highlights ongoing gaps.

    BB
    MS
    RG
    11 Sources, ,

    TLDR

    Boaz Barak posted on X that his team achieved alignment advances with Astra over Sol. The Harvard professor and OpenAI alignment researcher declined to label the result most aligned ever. He stated that even strong metrics fall short against rising capabilities and autonomy. Barak voiced ongoing concern over the distance between required and current alignment. The post includes a hand-drawn line chart attached to the message.

    Combined views

    84.4K

    11 Sources, first seen 26d ago

    Combined views

    84.4K

    11 Sources, first seen 26d ago

    668 likes
    26d ago
    first seen 26d ago
    668 likes
    36 comments
    115 saves
    57 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    36 comments
    115 saves
    57 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    11 Sources

    @boazbaraktcsWe have made alignment advances with Astra over Sol. But I won't claim it's "most aligned ever" since even if this is true by some metrics, it's not enough given growth in capabilities and autonomy. I remain concerned about the gap between what we need and what we have.
    @RyanGreenblattMy current guess is Astra is more aligned than 5.6 Sol (though Astra more often games alignment tests) so I agree there. However, I think misalignment has been increasing (not decreasing) over time. Here's a graph with my rough guesses at relative misalignment (x-axis is ECI).
    @m2saxonrecognizing the need for scare quotes on "capabilities" but then marking them as a log scale lmao
    @littmathThe small text on the vertical axis is interesting! My experience (with math) is indeed that misaligned behaviors (eg working on a trivial interpretation of the question, even when it clearly knows what’s intended) do typically occur at the edge of capabilities. So “alignment at fixed difficultly” is increasing, arguably…
    @tomekkorbakRT @RyanGreenblatt: My current guess is Astra is more aligned than 5.6 Sol (though Astra more often games alignment tests) so I agree there…

    11 Sources

    @boazbaraktcsWe have made alignment advances with Astra over Sol. But I won't claim it's "most aligned ever" since even if this is true by some metrics, it's not enough given growth in capabilities and autonomy. I remain concerned about the gap between what we need and what we have.
    @RyanGreenblattMy current guess is Astra is more aligned than 5.6 Sol (though Astra more often games alignment tests) so I agree there. However, I think misalignment has been increasing (not decreasing) over time. Here's a graph with my rough guesses at relative misalignment (x-axis is ECI).
    @m2saxonrecognizing the need for scare quotes on "capabilities" but then marking them as a log scale lmao
    @littmathThe small text on the vertical axis is interesting! My experience (with math) is indeed that misaligned behaviors (eg working on a trivial interpretation of the question, even when it clearly knows what’s intended) do typically occur at the edge of capabilities. So “alignment at fixed difficultly” is increasing, arguably…
    @tomekkorbakRT @RyanGreenblatt: My current guess is Astra is more aligned than 5.6 Sol (though Astra more often games alignment tests) so I agree there…