• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Dwarkesh announces an episode on automated AI research and what comes next

    The host’s agenda asks how automated AI researchers will be trained, whether long-horizon reinforcement learning will elicit artificial general intelligence, and how much progress is explained by data.

    MM
    RS
    DP
    18 Sources, ,

    TLDR

    Dwarkesh says he gathered AI researchers from what he calls “openish” companies to discuss the field’s next steps. Baseten says one of its own joins the episode. The host’s agenda covers automated AI research, reinforcement learning over longer tasks, Chinese labs’ progress and the role of data. It also includes arguments against recursive self-improvement, the gap between simulation and reality, and AI timelines.

    Combined views

    1.4M

    18 Sources, first seen 22d ago

    Combined views

    1.4M

    18 Sources, first seen 22d ago

    6.3K likes
    22d ago
    first seen 22d ago
    6.3K likes
    229 comments
    3.6K saves
    684 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    229 comments
    3.6K saves
    684 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    18 Sources

    @sebkrierRT @herbiebradley: Some takes about RSI from discussions with many smart researchers & thinkers: 1. Many RSI (or automated AI R&D) debates…
    @dwarkesh_spNew episode with @johnschulman2, @oneill_c and @BerenMillidge. I got together with some of the most insightful AI researchers I know who are at the openish companies, because I wanted to hear the details of what's actually happening at the frontier and what comes next. 0:00:00 – Steelmanning the case against RSI 0:18:39 – What’s driving the Chinese labs’ progress 0:28:06 – How will automated AI researchers be trained 0:33:51 – Will long-horizon RL elicit AGI? 0:45:24 – The sim-to-real gap 1:00:33 – How much progress is explained by data? 1:18:03 – Why is RL working so well? 1:24:54 – Move 37 and entropy collapse 1:28:31 – Rapid-fire timelines
    @oneill_cI think one of the more interesting things we debated here is whether RSI is a cumulative task. Attention plus MoE plus GRPO etc seems to me like a line in the sand that you can just add to the stack once you discover it. You don't need to take five steps back to take 10 steps forward. But a lot of the work in the world isn't this clean and it certainly isn't this stationary eg legal work. This leads to some perhaps unintuitive predictions such as why RSI might land before continual learning (and why it's going to be hard to get off the current paradigm even if it's wrong) Thanks for having me @dwarkesh_sp!
    @BerenMillidgeIt was a great conversation with @johnschulman2 and @oneill_c and I definitely learned a lot. Thanks @dwarkesh_sp for pulling this together! Understanding where we stand with RSI and how well current RL methods scale is an extremely important and interesting question
    @HesamationDwarkesh: When do we get AI that beats top human experts at all computer-based cognitive work? Beren: I would say 3–4 years. Dwarkesh: THE FUCK? 😭
    @DanielleFongRT @Hesamation: Dwarkesh: When do we get AI that beats top human experts at all computer-based cognitive work? Beren: I would say 3–4 year…
    @RichardSocherWe can safely advance science and society with AI or get distracted with scifi doom scenarios that require magical extrapolation. I really enjoyed this podcast on RSI for science. https://www.youtube.com/watch?v=kyyLku5F3bo
    @tessybartonRT @dwarkesh_sp: New episode with @johnschulman2, @oneill_c and @BerenMillidge. I got together with some of the most insightful AI researc…
    @basetenOur own @oneill_c joins @dwarkesh_sp, @johnschulman2, and @BerenMillidge for a deep discussion on long-horizon RL, automated AI research, and what comes next at the frontier.
    @thinkymachinesOur own @johnschulman2 talks with Dwarkesh about where human judgment still matters as models improve and self-improve: teaching them to handle messy real-world tasks, applying taste to what works in the long run, and, above all, specifying what we actually want.

    18 Sources

    @sebkrierRT @herbiebradley: Some takes about RSI from discussions with many smart researchers & thinkers: 1. Many RSI (or automated AI R&D) debates…
    @dwarkesh_spNew episode with @johnschulman2, @oneill_c and @BerenMillidge. I got together with some of the most insightful AI researchers I know who are at the openish companies, because I wanted to hear the details of what's actually happening at the frontier and what comes next. 0:00:00 – Steelmanning the case against RSI 0:18:39 – What’s driving the Chinese labs’ progress 0:28:06 – How will automated AI researchers be trained 0:33:51 – Will long-horizon RL elicit AGI? 0:45:24 – The sim-to-real gap 1:00:33 – How much progress is explained by data? 1:18:03 – Why is RL working so well? 1:24:54 – Move 37 and entropy collapse 1:28:31 – Rapid-fire timelines
    @oneill_cI think one of the more interesting things we debated here is whether RSI is a cumulative task. Attention plus MoE plus GRPO etc seems to me like a line in the sand that you can just add to the stack once you discover it. You don't need to take five steps back to take 10 steps forward. But a lot of the work in the world isn't this clean and it certainly isn't this stationary eg legal work. This leads to some perhaps unintuitive predictions such as why RSI might land before continual learning (and why it's going to be hard to get off the current paradigm even if it's wrong) Thanks for having me @dwarkesh_sp!
    @BerenMillidgeIt was a great conversation with @johnschulman2 and @oneill_c and I definitely learned a lot. Thanks @dwarkesh_sp for pulling this together! Understanding where we stand with RSI and how well current RL methods scale is an extremely important and interesting question
    @HesamationDwarkesh: When do we get AI that beats top human experts at all computer-based cognitive work? Beren: I would say 3–4 years. Dwarkesh: THE FUCK? 😭
    @DanielleFongRT @Hesamation: Dwarkesh: When do we get AI that beats top human experts at all computer-based cognitive work? Beren: I would say 3–4 year…
    @RichardSocherWe can safely advance science and society with AI or get distracted with scifi doom scenarios that require magical extrapolation. I really enjoyed this podcast on RSI for science. https://www.youtube.com/watch?v=kyyLku5F3bo
    @tessybartonRT @dwarkesh_sp: New episode with @johnschulman2, @oneill_c and @BerenMillidge. I got together with some of the most insightful AI researc…
    @basetenOur own @oneill_c joins @dwarkesh_sp, @johnschulman2, and @BerenMillidge for a deep discussion on long-horizon RL, automated AI research, and what comes next at the frontier.
    @thinkymachinesOur own @johnschulman2 talks with Dwarkesh about where human judgment still matters as models improve and self-improve: teaching them to handle messy real-world tasks, applying taste to what works in the long run, and, above all, specifying what we actually want.