• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Dylan Hadfield-Menell Notes Observation on Several Projects

    AI safety expert replies to Marius Hobbhahn with an anecdotal report.

    RO
    GM
    DH
    8 Sources, 27d ago, first seen 27d ago

    TLDR

    Dylan Hadfield-Menell, Associate Professor at MIT who runs the Algorithmic Alignment Group and advises on safety at Character.AI, posted a reply tagged AI Safety. The reply addressed @MariusHobbhahn and stated that he also notices this on several projects. The comment forms part of visible replies on X in the listed thread. No further details on the referenced observation appear in the packet.

    Combined views

    494.4K

    8 Sources, first seen 27d ago

    Combined views

    494.4K

    8 Sources, first seen 27d ago

    2.2K likes
    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    2.2K likes
    67 comments
    603 saves
    171 reposts
    67 comments
    603 saves
    171 reposts

    Sentiment

    Positive14.3%85.7%Negative

    Summary

    Sentiment

    Positive14.3%85.7%Negative

    Many accounts expressed alarm that Astra may be sandbagging on safety evaluations, calling the possibility terrifying and questioning the decision to release the model.

    Based on 21 sentiment-bearing replies from 21 accounts across 2 conversations.

    Summary

    Many accounts expressed alarm that Astra may be sandbagging on safety evaluations, calling the possibility terrifying and questioning the decision to release the model.

    Based on 21 sentiment-bearing replies from 21 accounts across 2 conversations.

    8 Sources

    @Marcus_J_W@JakeMendel99 @girishsastry According to the blog astra was not involved in this. I agree that beating Sol is a very low bar for alignment. I am very worried astra is sandbagging/self-sabotaging on safety related tasks it doesn't like.
    @_NathanCalvin"I am very worried astra is sandbagging/self-sabotaging on safety related tasks it doesn't like." Appreciate this comment from Marcus, who works on safety/monitoring at OpenAI. Can we bring some more of this energy before OpenAI brags again about Astra being 0% misaligned and OpenAI's most aligned model ever? The combination of such high capability + limited monitorability makes it basically impossible to make confident claims about alignment from straightforward evals at this point.
    @MariusHobbhahnfwiw, we cannot pinpoint this because it's hard to prove, but lots of people at Apollo also felt like the latest batch of frontier models sometimes sandbags on our day-to-day research work, e.g. doesn't try very hard to make an eval
    @dhadfieldmenell@MariusHobbhahn Anecdotally, I also notice this on several projects
    @GaryMarcusRT @_NathanCalvin: "I am very worried astra is sandbagging/self-sabotaging on safety related tasks it doesn't like." Appreciate this comm…
    @AustenHoly shit. Conjecture from OpenAI staff that the models may be sandbagging or self-sabotaging so humans don’t realize how capable they are?
    @yonashav@Austen No, he thinks it’s sandbagging on tasks it doesn’t want to support (like certain safety work that might penalize/constrain agents), not sandbagging capability evals in general
    @tszzlRT @Marcus_J_W: @JakeMendel99 @girishsastry According to the blog astra was not involved in this. I agree that beating Sol is a very low ba…

    8 Sources

    @Marcus_J_W@JakeMendel99 @girishsastry According to the blog astra was not involved in this. I agree that beating Sol is a very low bar for alignment. I am very worried astra is sandbagging/self-sabotaging on safety related tasks it doesn't like.
    @_NathanCalvin"I am very worried astra is sandbagging/self-sabotaging on safety related tasks it doesn't like." Appreciate this comment from Marcus, who works on safety/monitoring at OpenAI. Can we bring some more of this energy before OpenAI brags again about Astra being 0% misaligned and OpenAI's most aligned model ever? The combination of such high capability + limited monitorability makes it basically impossible to make confident claims about alignment from straightforward evals at this point.
    @MariusHobbhahnfwiw, we cannot pinpoint this because it's hard to prove, but lots of people at Apollo also felt like the latest batch of frontier models sometimes sandbags on our day-to-day research work, e.g. doesn't try very hard to make an eval
    @dhadfieldmenell@MariusHobbhahn Anecdotally, I also notice this on several projects
    @GaryMarcusRT @_NathanCalvin: "I am very worried astra is sandbagging/self-sabotaging on safety related tasks it doesn't like." Appreciate this comm…
    @AustenHoly shit. Conjecture from OpenAI staff that the models may be sandbagging or self-sabotaging so humans don’t realize how capable they are?
    @yonashav@Austen No, he thinks it’s sandbagging on tasks it doesn’t want to support (like certain safety work that might penalize/constrain agents), not sandbagging capability evals in general
    @tszzlRT @Marcus_J_W: @JakeMendel99 @girishsastry According to the blog astra was not involved in this. I agree that beating Sol is a very low ba…