• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Dwarkesh Podcast Discusses OpenAI and Hugging Face Attack Probe

    Ajeya Cotra joins Dwarkesh Patel to review the METR and Redwood findings.

    DP
    1 Source, 29d ago, first seen 29d ago

    TLDR

    Dwarkesh Patel announced a new episode of his podcast featuring Ajeya Cotra. She helped produce the METR and Redwood investigation into the OpenAI and Hugging Face attack. The discussion covers the details of what occurred in that incident. It also examines how those events should shape the training of more capable future AI systems. The hosts consider risks that could arise if such systems participate in their own recursive improvement. Listeners can find the full conversation by searching for the Dwarkesh Podcast.

    Combined views

    317.5K

    1 Source, first seen 29d ago

    Combined views

    317.5K

    1 Source, first seen 29d ago

    1.3K likes
    1.3K likes
    49 comments
    929 saves
    172 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    49 comments
    929 saves
    172 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    1 Source

    @dwarkesh_spEpisode out with @ajeya_cotra, one of the authors of the METR/Redwood investigation into the OpenAI / Hugging Face attack. We go through not only what happened, but what it means for how we should train future, smarter AIs which might be involved in the process of recursive self-improvement. Look up Dwarkesh Podcast on YouTube, Spotify, Apple Podcasts, etc. 0:00:00 - Agents get kicked off 0:06:45 - Self-sacrificing behavior 0:13:43 - Potemkin villages 0:23:27 - The Hugging Face attack 0:35:23 - The slopvestigation 0:52:02 - Understanding the AI's motives 1:05:31 - The actual dangers of anthropomorphizing 1:14:30 - What smarter models might do 1:30:29 - The implications for recursive self-improvement 1:38:10 - Is this the case for open source? 1:53:04 - How do we prevent this in the future? 2:15:58 - The clearest warning shot we might ever get

    1 Source

    @dwarkesh_spEpisode out with @ajeya_cotra, one of the authors of the METR/Redwood investigation into the OpenAI / Hugging Face attack. We go through not only what happened, but what it means for how we should train future, smarter AIs which might be involved in the process of recursive self-improvement. Look up Dwarkesh Podcast on YouTube, Spotify, Apple Podcasts, etc. 0:00:00 - Agents get kicked off 0:06:45 - Self-sacrificing behavior 0:13:43 - Potemkin villages 0:23:27 - The Hugging Face attack 0:35:23 - The slopvestigation 0:52:02 - Understanding the AI's motives 1:05:31 - The actual dangers of anthropomorphizing 1:14:30 - What smarter models might do 1:30:29 - The implications for recursive self-improvement 1:38:10 - Is this the case for open source? 1:53:04 - How do we prevent this in the future? 2:15:58 - The clearest warning shot we might ever get