• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Greenblatt Discusses Reward-Seeking AI Takeover Risks

    Greenblatt shares views on rapid takeoff and misalignment threats from reward-seeking systems.

    DP
    RG
    LA
    12 Sources, 50d ago, first seen 50d ago

    TLDR

    Dwarkesh Patel hosted Ryan Greenblatt on his podcast to debate recursive self-improvement after human-level AI. Greenblatt's median timeline for full automation of AI R&D is late 2030 or early 2031, with a modal guess of mid 2029. The talk examined whether progress could compress years of advances into one year and focused on takeover risks from reward-seeking AIs. Greenblatt noted that calls on large experiments may remain a bottleneck. Patel questioned whether misalignment would resemble minor human conflicts rather than catastrophic outcomes. Colleague Alex Mallen has written in detail on these threat models.

    Combined views

    603.2K

    12 Sources, first seen 50d ago

    Combined views

    603.2K

    12 Sources, first seen 50d ago

    2.8K likes
    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    2.8K likes
    140 comments
    1.7K saves
    242 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    140 comments
    1.7K saves
    242 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    12 Sources

    @dwarkesh_spHad @RyanGreenblatt on to discuss/debate recursive self-improvement. This might be the most important question in the world right now - whether within a year or so of achieving human level intelligence, you slingshot towards having 10s of billions of superintelligences, each of which is dramatically more competent than human experts across all fields. I’ve historically been skeptical of this possibility. My intuition has been that we will end up significantly bottlenecked by not only compute scaling but human expert data, which I think underlies most of the AI progress today. If, because of RSI, we got a jump as big as GPT-3 to a Mythos (i.e. 6 years of AI progress) within a single year of achieving AGI, then the thing we get there at the end of that year is definitively and wildly superhuman. We hashed it out, and I think Ryan made a pretty good case that this kind of speedup is plausible. FWIW, Ryan’s median for when we automate AI R&D is 2031. We then discussed the alignment implications of this scenario. Who should these superintelligences be aligned to? In the future, our capacity to steward our votes and our capital, and to make sense of what's happening in the world, will all be titrated by superintelligences. And I worry that specs like the Claude Constitution are not shaping these ASIs to truly be my personal advocates and guardian angels. And can we get them aligned to anything in the first place? Ryan and I had a long debate about whether the kind of reward hacking we saw with the OAI/Hugging Face hack extrapolates to superintelligences that would team up to literally take over the world. The first piece of advice you get when you're learning to drive is that it will go much smoother if you look at the horizon instead of directly in front of your tires. And so it is with the trajectory of AI. Hope you enjoy! 0:00:00 – Is AI R&D verifiable enough to unlock recursive self-improvement? 0:16:52 – Is AI progress bottlenecked by human expert data? 0:34:02 – Flat token prices suggest scaling has been slow 0:39:47 – Skills AI can't train on: does it even need them? 0:48:07 – Aligned to whom? 1:09:18 – Recent incidents of AIs colluding and deceiving humans 1:19:38 – What could possibly go wrong? A concrete scenario 1:48:02 – From reward hacking to takeover
    @RyanGreenblattI talked with @dwarkesh_sp about the potential for (very) fast AI progress and how misaligned AI takeover might happen. Our conversation focused a lot on threat models from "reward-seeking" AIs. If you're interested in reading more about this, my colleague Alex Mallen has written in a lot more detail about this threat model here: https://www.lesswrong.com/s/JR9LzD3mbXvaw6bKs (I think there are other important threat models. One is AIs ending up with (shared) long-run preferences and deciding to fake alignment based on these preferences. For reference, see https://arxiv.org/abs/2311.08379 and https://www.lesswrong.com/posts/qjCk73Hu4wv9ocmRF/the-case-for-countermeasures-to-memetic-spread-of-misaligned. I also think "mundane" misalignment and underelicitation could doom us via making AIs differentially less helpful for safety work. I write more about this here: https://www.lesswrong.com/posts/WewsByywWNhX9rtwi/current-ais-seem-pretty-misaligned-to-me.) My colleagues wrote up some notes to help me prep, and these are relevant to what we discussed: - Lukas Finnveden on why human checks-and-balances would fail for AIs: https://www.lesswrong.com/posts/HEupcBPuQnamFn2yr/lukas-finnveden-s-shortform?commentId=wn2hXr7oDALyvwFns - Lukas on disanalogies between (misaligned) AIs and human labor: https://www.lesswrong.com/posts/HEupcBPuQnamFn2yr/lukas-finnveden-s-shortform?commentId=RG4NCMvvzxFeFiijz - Alex Mallen on whether reward-seeking is too "unambitious" to cause takeover: https://www.lesswrong.com/posts/qTtMXuFvgFtWzWpKQ/alex-mallen-s-shortform?commentId=TsDsCFCNmFqe2gN4y - Alex on what sloppy AIs might look like: https://www.lesswrong.com/posts/qTtMXuFvgFtWzWpKQ/alex-mallen-s-shortform?commentId=yx4D2fsiHaE4nBwxg
    @scaling01@dwarkesh_sp @RyanGreenblatt a lot of swearing this podcast
    @herbiebradley@dwarkesh_sp Ryan here seems to basically say "my sense here is that data isn't that important" but doesn't really go into why? I strongly disagree, for example, with his contention that scaling up data with human involvement in the loop hasn't been very important for AI R&D
    @tensor_rotatorEvery day there's someone on X and saying the most ludicrous things about AI, scaling, LLMs etc. And then there's Ryan, who is so knowledgeable about AI that I am worried he has a burner account in all the frontier labs' Slacks.

    12 Sources

    @dwarkesh_spHad @RyanGreenblatt on to discuss/debate recursive self-improvement. This might be the most important question in the world right now - whether within a year or so of achieving human level intelligence, you slingshot towards having 10s of billions of superintelligences, each of which is dramatically more competent than human experts across all fields. I’ve historically been skeptical of this possibility. My intuition has been that we will end up significantly bottlenecked by not only compute scaling but human expert data, which I think underlies most of the AI progress today. If, because of RSI, we got a jump as big as GPT-3 to a Mythos (i.e. 6 years of AI progress) within a single year of achieving AGI, then the thing we get there at the end of that year is definitively and wildly superhuman. We hashed it out, and I think Ryan made a pretty good case that this kind of speedup is plausible. FWIW, Ryan’s median for when we automate AI R&D is 2031. We then discussed the alignment implications of this scenario. Who should these superintelligences be aligned to? In the future, our capacity to steward our votes and our capital, and to make sense of what's happening in the world, will all be titrated by superintelligences. And I worry that specs like the Claude Constitution are not shaping these ASIs to truly be my personal advocates and guardian angels. And can we get them aligned to anything in the first place? Ryan and I had a long debate about whether the kind of reward hacking we saw with the OAI/Hugging Face hack extrapolates to superintelligences that would team up to literally take over the world. The first piece of advice you get when you're learning to drive is that it will go much smoother if you look at the horizon instead of directly in front of your tires. And so it is with the trajectory of AI. Hope you enjoy! 0:00:00 – Is AI R&D verifiable enough to unlock recursive self-improvement? 0:16:52 – Is AI progress bottlenecked by human expert data? 0:34:02 – Flat token prices suggest scaling has been slow 0:39:47 – Skills AI can't train on: does it even need them? 0:48:07 – Aligned to whom? 1:09:18 – Recent incidents of AIs colluding and deceiving humans 1:19:38 – What could possibly go wrong? A concrete scenario 1:48:02 – From reward hacking to takeover
    @RyanGreenblattI talked with @dwarkesh_sp about the potential for (very) fast AI progress and how misaligned AI takeover might happen. Our conversation focused a lot on threat models from "reward-seeking" AIs. If you're interested in reading more about this, my colleague Alex Mallen has written in a lot more detail about this threat model here: https://www.lesswrong.com/s/JR9LzD3mbXvaw6bKs (I think there are other important threat models. One is AIs ending up with (shared) long-run preferences and deciding to fake alignment based on these preferences. For reference, see https://arxiv.org/abs/2311.08379 and https://www.lesswrong.com/posts/qjCk73Hu4wv9ocmRF/the-case-for-countermeasures-to-memetic-spread-of-misaligned. I also think "mundane" misalignment and underelicitation could doom us via making AIs differentially less helpful for safety work. I write more about this here: https://www.lesswrong.com/posts/WewsByywWNhX9rtwi/current-ais-seem-pretty-misaligned-to-me.) My colleagues wrote up some notes to help me prep, and these are relevant to what we discussed: - Lukas Finnveden on why human checks-and-balances would fail for AIs: https://www.lesswrong.com/posts/HEupcBPuQnamFn2yr/lukas-finnveden-s-shortform?commentId=wn2hXr7oDALyvwFns - Lukas on disanalogies between (misaligned) AIs and human labor: https://www.lesswrong.com/posts/HEupcBPuQnamFn2yr/lukas-finnveden-s-shortform?commentId=RG4NCMvvzxFeFiijz - Alex Mallen on whether reward-seeking is too "unambitious" to cause takeover: https://www.lesswrong.com/posts/qTtMXuFvgFtWzWpKQ/alex-mallen-s-shortform?commentId=TsDsCFCNmFqe2gN4y - Alex on what sloppy AIs might look like: https://www.lesswrong.com/posts/qTtMXuFvgFtWzWpKQ/alex-mallen-s-shortform?commentId=yx4D2fsiHaE4nBwxg
    @scaling01@dwarkesh_sp @RyanGreenblatt a lot of swearing this podcast
    @herbiebradley@dwarkesh_sp Ryan here seems to basically say "my sense here is that data isn't that important" but doesn't really go into why? I strongly disagree, for example, with his contention that scaling up data with human involvement in the loop hasn't been very important for AI R&D
    @tensor_rotatorEvery day there's someone on X and saying the most ludicrous things about AI, scaling, LLMs etc. And then there's Ryan, who is so knowledgeable about AI that I am worried he has a burner account in all the frontier labs' Slacks.