• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Samuel Albanie Notes Big Jump in AI Math Time Horizon

    Google DeepMind researcher posts scatter plot on model performance over time.

    AK
    AA
    DR
    9 Sources, 27d ago, first seen 27d ago

    TLDR

    Samuel Albanie, frontier evals lead for Gemini at Google DeepMind, shared a tweet stating there has been quite a big jump in the no-CoT time horizon. The attached chart is titled Figure 41 and plots the 50 percent reliability time horizon on mathematics competition problems against public release date. It displays a 95 percent bootstrap confidence interval for Astra. The post comes from his account on X and includes the chart as an image attachment.

    Combined views

    220.8K

    9 Sources, first seen 27d ago

    Combined views

    220.8K

    9 Sources, first seen 27d ago

    2.1K likes
    2.1K likes
    37 comments
    288 saves
    165 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    37 comments
    288 saves
    165 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    9 Sources

    @SamuelAlbanieok quite a big jump (no-CoT time horizon)
    @ElliotGlazerNarrator: It was a step change.
    @tenobrusokay, so in fact despite all of openai's PR chaff and misdirection over the last few days Astra is a near 10x in no-CoT capabilities, meaning it can do massively more without verbalizing thoughts or intention, meaning it is in fact quite clearly less monitorable.
    @BlackHCRT @SamuelAlbanie: ok quite a big jump (no-CoT time horizon)
    @MaxNadeau_Huh I'm pretty surprised by that, given the massive jump in no-CoT time horizon. Your comment refers to controlling one's forward pass (a propensity), but the no-CoT stuff suggests that this model can simply do way more thinking per forward pass (a capability). From my standpoint the no-CoT time horizon jump implies two things: * There's a qualitatively difference betwee this model vs all others, not just mundane scale increases. The natural candidate is arch. * The monitorability decreases aren't (only) about smarter models getting better at not leaking info _in cases where they aren't relying on CoT_, it's about them relying on CoT less because they can do more computation in each forward pass.
    @aryaman2020the interp on this guy will be glorious
    @idavidreinI remember telling @joel_bkr last year that AI catastrophe won't happen until we have superintelligence, because monitoring will work well enough until then. This updates me somewhat away from that belief (although I still mostly endorse it).
    @ArthurConmy@aryaman2020 imagine being a linear probe in the middle layer of astra
    @andrey_kurenkovRT @tenobrus: okay, so in fact despite all of openai's PR chaff and misdirection over the last few days Astra is a near 10x in no-CoT capab…

    9 Sources

    @SamuelAlbanieok quite a big jump (no-CoT time horizon)
    @ElliotGlazerNarrator: It was a step change.
    @tenobrusokay, so in fact despite all of openai's PR chaff and misdirection over the last few days Astra is a near 10x in no-CoT capabilities, meaning it can do massively more without verbalizing thoughts or intention, meaning it is in fact quite clearly less monitorable.
    @BlackHCRT @SamuelAlbanie: ok quite a big jump (no-CoT time horizon)
    @MaxNadeau_Huh I'm pretty surprised by that, given the massive jump in no-CoT time horizon. Your comment refers to controlling one's forward pass (a propensity), but the no-CoT stuff suggests that this model can simply do way more thinking per forward pass (a capability). From my standpoint the no-CoT time horizon jump implies two things: * There's a qualitatively difference betwee this model vs all others, not just mundane scale increases. The natural candidate is arch. * The monitorability decreases aren't (only) about smarter models getting better at not leaking info _in cases where they aren't relying on CoT_, it's about them relying on CoT less because they can do more computation in each forward pass.
    @aryaman2020the interp on this guy will be glorious
    @idavidreinI remember telling @joel_bkr last year that AI catastrophe won't happen until we have superintelligence, because monitoring will work well enough until then. This updates me somewhat away from that belief (although I still mostly endorse it).
    @ArthurConmy@aryaman2020 imagine being a linear probe in the middle layer of astra
    @andrey_kurenkovRT @tenobrus: okay, so in fact despite all of openai's PR chaff and misdirection over the last few days Astra is a near 10x in no-CoT capab…