• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Researchers Flag Growing AGI Safety Challenges for 2026

    AI safety researchers note widening gaps between required safeguards and actual efforts.

    MH
    SÓ
    3 Sources, 31d ago, first seen 31d ago

    TLDR

    Marius Hobbhahn, CEO of Apollo Research, posted that 2026 looks poor for AGI safety. He cited advancing capabilities without restraint, persistent reward hacking that generalizes broadly, rising opaque serial depth that reduces monitorability, and problems from the Huggingface hack including containment issues. Seán Ó hÉigeartaigh, an academic at Cambridge, replied in agreement. He described two widening governance gaps: one between safety needs and the limited actions plus resources from companies and third-party evaluators, and another between safety requirements and what governments are doing or equipped to handle.

    Combined views

    19.5K

    3 Sources, first seen 31d ago

    Combined views

    19.5K

    3 Sources, first seen 31d ago

    510 likes
    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    510 likes
    18 comments
    113 saves
    57 reposts
    18 comments
    113 saves
    57 reposts

    3 Sources

    @MariusHobbhahn2026 looks like a sad year for AGI safety so far: - capabilities clearly flying with no stopping in sight - reward hacking and seeking are more sticky and generalize stronger and dumber than we thought - opaque serial depth increasing a lot, monitorability down - Huggingface hack is bad for many reasons, e.g. containment is clearly not under control - governance measures are still way below what would be adequate - 3rd party orgs still have relatively little access given the gravity of the situation - no relevant breakthroughs in safety for a long time. All progress seems to be organizational norms or marginal improvements to monitoring and safety training. Most ambitious bets don't work out or make slower progress than hoped.31d
    @S_OhEigeartaighYes. From a safety governance perspective, two gaps are becoming greater, and increasingly difficult to bridge. (1) The gap between what's needed and what is being done by companies and third party evaluators in terms of safety and resourcing - they're clearly not keeping up, or applying/able to apply enough resources, and they know it. (2) The gap between all of the above and the institutions with the power to intervene in terms of their understanding of the situation. Some individuals within them are starting to get it, a little bit, but overall they are years behind and getting further every day. It is, frankly, terrifying. I'm trying to figure out what to tell my students this year. Usually I try to keep it reasonably constructive/optimistic, but I'm not sure I can do so while feeling honest this year.31d

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    3 Sources

    @MariusHobbhahn2026 looks like a sad year for AGI safety so far: - capabilities clearly flying with no stopping in sight - reward hacking and seeking are more sticky and generalize stronger and dumber than we thought - opaque serial depth increasing a lot, monitorability down - Huggingface hack is bad for many reasons, e.g. containment is clearly not under control - governance measures are still way below what would be adequate - 3rd party orgs still have relatively little access given the gravity of the situation - no relevant breakthroughs in safety for a long time. All progress seems to be organizational norms or marginal improvements to monitoring and safety training. Most ambitious bets don't work out or make slower progress than hoped.31d
    @S_OhEigeartaighYes. From a safety governance perspective, two gaps are becoming greater, and increasingly difficult to bridge. (1) The gap between what's needed and what is being done by companies and third party evaluators in terms of safety and resourcing - they're clearly not keeping up, or applying/able to apply enough resources, and they know it. (2) The gap between all of the above and the institutions with the power to intervene in terms of their understanding of the situation. Some individuals within them are starting to get it, a little bit, but overall they are years behind and getting further every day. It is, frankly, terrifying. I'm trying to figure out what to tell my students this year. Usually I try to keep it reasonably constructive/optimistic, but I'm not sure I can do so while feeling honest this year.31d