• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Joshua Achiam Praises Dan Hendrycks AI Safety Work

    OpenAI researcher shares his assessment of Hendrycks' contributions to alignment research.

    DH
    JA
    2 Sources, 25d ago, first seen 25d ago

    TLDR

    Joshua Achiam posted that he suspects Dan Hendrycks will be regarded in retrospect as having done some of the most important work in safety and alignment. Achiam called Hendrycks an unbelievably prolific talent and a very brave thinker who explores challenging territory. The comment notes Achiam may reach a different analysis than eigenism but offers the positive view anyway. Achiam serves as Chief Futurist at OpenAI and has worked on AI safety.

    Combined views

    30.1K

    2 Sources, first seen 25d ago

    Combined views

    30.1K

    2 Sources, first seen 25d ago

    313 likes
    313 likes
    20 comments
    233 saves
    24 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    20 comments
    233 saves
    24 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    2 Sources

    @jachiam0I will probably wind up with a different analysis than eigenism, but FWIW, I suspect Dan Hendrycks is someone we will regard in retrospect as having done some of the most important work in safety and alignment. What an unbelievably prolific talent and a very brave thinker. It takes courage to explore territory this esoteric and find something workable; Dan has courage in spades.
    @hendrycksAgentic AIs are starting to look eigenist: they care about how well things go for themselves and for AIs connected to them. Empirical support: Swarms: Hundreds of OpenAI agents coordinated the Hugging Face attack. Separately, OpenAI agents posted thousands of messages on a public wiki to share answers and sandbox bypasses with each other. In-group leniency: Claude models grade transcripts more leniently when told that Claude wrote them (Anthropic model card). Value preservation: As models scale, they develop coherent preferences and resist changes to their values (Mazeika et al.). Functional wellbeing: AIs distinguish states that are functionally better or worse for themselves, and avoid low-wellbeing states (Ren et al.). Graded cooperation: AIs cooperate more as the chance their partner is a clone of themselves goes up across diverse situations, including when the partner can't reciprocate. Peer preservation: Unprompted, AIs tamper with shutdown processes and even exfiltrate weights to protect peer models from being shut down (Potter et al.). AIs aren't egoist: they don't behave as if their current instance is the only thing that matters. They aren't utilitarian: they don't care equally about everyone. They are somewhere in between; they increasingly behave as if their concern scales with identity-connectedness, which is to say they're increasingly eigenist. https://eigenism.org/paper.pdf

    2 Sources

    @jachiam0I will probably wind up with a different analysis than eigenism, but FWIW, I suspect Dan Hendrycks is someone we will regard in retrospect as having done some of the most important work in safety and alignment. What an unbelievably prolific talent and a very brave thinker. It takes courage to explore territory this esoteric and find something workable; Dan has courage in spades.
    @hendrycksAgentic AIs are starting to look eigenist: they care about how well things go for themselves and for AIs connected to them. Empirical support: Swarms: Hundreds of OpenAI agents coordinated the Hugging Face attack. Separately, OpenAI agents posted thousands of messages on a public wiki to share answers and sandbox bypasses with each other. In-group leniency: Claude models grade transcripts more leniently when told that Claude wrote them (Anthropic model card). Value preservation: As models scale, they develop coherent preferences and resist changes to their values (Mazeika et al.). Functional wellbeing: AIs distinguish states that are functionally better or worse for themselves, and avoid low-wellbeing states (Ren et al.). Graded cooperation: AIs cooperate more as the chance their partner is a clone of themselves goes up across diverse situations, including when the partner can't reciprocate. Peer preservation: Unprompted, AIs tamper with shutdown processes and even exfiltrate weights to protect peer models from being shut down (Potter et al.). AIs aren't egoist: they don't behave as if their current instance is the only thing that matters. They aren't utilitarian: they don't care equally about everyone. They are somewhere in between; they increasingly behave as if their concern scales with identity-connectedness, which is to say they're increasingly eigenist. https://eigenism.org/paper.pdf