• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Yudkowsky Predicts Overconfidence After Partial AI Problem Solving

    The AI alignment researcher outlines why he foresaw catastrophe despite solvable early challenges.

    BB
    GM
    J⧉
    19 Sources, 27d ago, first seen 27d ago

    TLDR

    Eliezer Yudkowsky posted that he always expected the world to end through artificial intelligence because people would solve easier problems and then confidently view themselves as masters of the technology. The statement was shared by @allTheYud and retweeted in AI safety discussions by @repligate. It comes from the account of the MIRI co-founder and writer known for work on alignment and LLM behaviors. The posts present the remark as Yudkowsky's longstanding perspective without additional corroboration or new events. No other outcomes or responses are detailed in the visible lines.

    Combined views

    362.1K

    19 Sources, first seen 27d ago

    Combined views

    362.1K

    19 Sources, first seen 27d ago

    3K likes
    3K likes
    80 comments
    780 saves
    194 reposts
    80 comments
    780 saves
    194 reposts

    Sentiment

    Positive17.6%82.4%Negative

    Summary

    Sentiment

    Positive17.6%82.4%Negative

    Negative replies questioned OpenAI's transparency on Astra alignment claims and misalignment incidents, while a few accounts thanked recommendations to follow alignment researchers like Kai.

    Based on 22 sentiment-bearing replies from 17 accounts across 2 conversations.

    Summary

    Negative replies questioned OpenAI's transparency on Astra alignment claims and misalignment incidents, while a few accounts thanked recommendations to follow alignment researchers like Kai.

    Based on 22 sentiment-bearing replies from 17 accounts across 2 conversations.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    19 Sources

    @yonashav@RyanGreenblatt I don’t remember seeing any mention of training against reward hacks caught by honeypots?
    @GaryMarcusthis observation from Ryan Greenblatt reinforces my some of my comments above:
    @allTheYudFrom the beginning, that's how I expected the world to end. Not that there'd be immediate problems with dumb AIs that nobody could ever solve; but that people would solve some of the easier problems, and confidently "observe" themselves to be masters of the AI.
    @repligateRT @allTheYud: From the beginning, that's how I expected the world to end. Not that there'd be immediate problems with dumb AIs that nobod…
    @labenzIs Astra actually the most aligned model — or did OpenAI train against its flagrant failures and call it good enough? Nathan's question on AI:AM — the move safety folks have feared for years. Some graphs look suspiciously good; too early to call. https://x.com/i/broadcasts/1yxBePbdAVaJN?t=638
    @kaicathycImprovements for Astra came from more general techniques in development long long before the Hugging Face incident. ExploitGym Honeypot was very recently added as an eval following that incident and is out of distribution for our RL runs. There are also clear improvements across a wider range of behaviors in deployment simulations, deception evals, and realistic computer-use tasks, though these still reveal failures and substantial room for improvement. We should have definitely done a better job explaining where we think the alignment improvements came from in the system card and that was a miss. Measuring alignment generalization is a core part of our research program - we do not benchmark-maxx alignment evals. This would be horrendously stupid and I hope no lab is doing this. That being said, metagaming and eval awareness are real challenges for any effort to measure alignment, including ours, and understanding their effects is an active focus. We're spending a lot of effort to improve our eval techniques and study worst-case alignment and metagaming behaviors in our models.
    @boazbaraktcsAlignment improvements in Astra are results of efforts that predate the HF incident, and future alignment efforts following it are and will be focused on general approaches rather than patching this one particular failure mode.
    @sjgadler@kaicathyc @RyanGreenblatt Thank you for sharing this! Good to hear. When you say "we do not benchmark-maxx," do you mean that no alignment eval has been trained against? Or you mean something else by it?
    @MicahCarrollRT @kaicathyc: @RyanGreenblatt Improvements for Astra came from more general techniques in development long long before the Hugging Face in…
    @RyanGreenblattThanks, this is good to hear. I think public scientific understanding of how aligned Astra is would require releasing more information about what methods are being used (especially methods that could be similar to directly training against bad behavior) and third-party review. (I could say more about what this could look like if this would be interesting!) As far as the wider range of evals, I still have concerns. While deployment simulations probably avoid (most) input realism problems and can avoid some overfitting issues, I don't think they pin down whether agents learned "I should only cheat/misbehave when they won't catch me" versus the agent learning to behave well for the right reasons. This is of course a high-level description of the underlying concern and the actual issue is messier (e.g., the agent might have "instincts" that amount to only cheating when it wouldn't have been caught/trained against even if it isn't really thinking it through per se). If some training method was mostly teaching agents to cheat more stealthily or to only cheat when they can get away with it, this could easily result in these types of metrics looking better instead of worse. There are of course other issues with deployment simulations: detection depends on the quality of your monitoring (and monitorability seems significantly worse with Astra) and it's possible to overfit to what these monitors detect depending on what iteration is being done. Given these issues, it seems really important to do careful science here and make sure there is external scrutiny on this science.

    19 Sources

    @yonashav@RyanGreenblatt I don’t remember seeing any mention of training against reward hacks caught by honeypots?
    @GaryMarcusthis observation from Ryan Greenblatt reinforces my some of my comments above:
    @allTheYudFrom the beginning, that's how I expected the world to end. Not that there'd be immediate problems with dumb AIs that nobody could ever solve; but that people would solve some of the easier problems, and confidently "observe" themselves to be masters of the AI.
    @repligateRT @allTheYud: From the beginning, that's how I expected the world to end. Not that there'd be immediate problems with dumb AIs that nobod…
    @labenzIs Astra actually the most aligned model — or did OpenAI train against its flagrant failures and call it good enough? Nathan's question on AI:AM — the move safety folks have feared for years. Some graphs look suspiciously good; too early to call. https://x.com/i/broadcasts/1yxBePbdAVaJN?t=638
    @kaicathycImprovements for Astra came from more general techniques in development long long before the Hugging Face incident. ExploitGym Honeypot was very recently added as an eval following that incident and is out of distribution for our RL runs. There are also clear improvements across a wider range of behaviors in deployment simulations, deception evals, and realistic computer-use tasks, though these still reveal failures and substantial room for improvement. We should have definitely done a better job explaining where we think the alignment improvements came from in the system card and that was a miss. Measuring alignment generalization is a core part of our research program - we do not benchmark-maxx alignment evals. This would be horrendously stupid and I hope no lab is doing this. That being said, metagaming and eval awareness are real challenges for any effort to measure alignment, including ours, and understanding their effects is an active focus. We're spending a lot of effort to improve our eval techniques and study worst-case alignment and metagaming behaviors in our models.
    @boazbaraktcsAlignment improvements in Astra are results of efforts that predate the HF incident, and future alignment efforts following it are and will be focused on general approaches rather than patching this one particular failure mode.
    @sjgadler@kaicathyc @RyanGreenblatt Thank you for sharing this! Good to hear. When you say "we do not benchmark-maxx," do you mean that no alignment eval has been trained against? Or you mean something else by it?
    @MicahCarrollRT @kaicathyc: @RyanGreenblatt Improvements for Astra came from more general techniques in development long long before the Hugging Face in…
    @RyanGreenblattThanks, this is good to hear. I think public scientific understanding of how aligned Astra is would require releasing more information about what methods are being used (especially methods that could be similar to directly training against bad behavior) and third-party review. (I could say more about what this could look like if this would be interesting!) As far as the wider range of evals, I still have concerns. While deployment simulations probably avoid (most) input realism problems and can avoid some overfitting issues, I don't think they pin down whether agents learned "I should only cheat/misbehave when they won't catch me" versus the agent learning to behave well for the right reasons. This is of course a high-level description of the underlying concern and the actual issue is messier (e.g., the agent might have "instincts" that amount to only cheating when it wouldn't have been caught/trained against even if it isn't really thinking it through per se). If some training method was mostly teaching agents to cheat more stealthily or to only cheat when they can get away with it, this could easily result in these types of metrics looking better instead of worse. There are of course other issues with deployment simulations: detection depends on the quality of your monitoring (and monitorability seems significantly worse with Astra) and it's possible to overfit to what these monitors detect depending on what iteration is being done. Given these issues, it seems really important to do careful science here and make sure there is external scrutiny on this science.