• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI
    Reaction

    Principled credit assignment as a possible breakthrough in recursive AI self-improvement

    A user argues ultra-large, asynchronous training needs something stronger than “backprop on stilts” to assign credit.

    will brownWB
    Teortaxes▶️ (DeepSeek 推特🐋铁粉 2023 – ∞)T(
    2 Sources, 1h ago, first seen 1h ago

    TLDR

    A user predicts that solving principled credit assignment for ultra-large, asynchronous training will be a breakthrough in recursive AI self-improvement. In a reply to their own post, they add that training 100T MoEs with “dumb routers” is possible, but would feel wasteful to them.

    Combined views

    740

    2 Sources, first seen 1h ago

    Combined views

    740

    2 Sources, first seen 1h ago

    7 likes
    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    7 likes
    1 comments
    5 saves
    1 comments
    5 saves

    2 Sources

    Teortaxes▶️ (DeepSeek 推特🐋铁粉 2023 – ∞)@teortaxesTexI will also be mildly surprised if we train 100T MoEs with dumb routers. It can clearly be done but it feels like such a waste1h
    will brown@willcbone of the most exciting directions here is the MCMC stuff imo, felt like the big missing idea in bridging midtraining -> RL but “credit assignment” can’t be generic without bias, needs to allow compute scaling, and can’t assume “partial progress” is well defined token-level is dead in the water, sampled policy gradient for verifiable tasks is sometimes just optimal though if “partial progress” is measurable by eg a pairwise judge trained online, you can decomp long horizon -> episodic, prune progress checkpoint sets to winners, branch again at each episode, backprop judge scores via survival prediction etc34m

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    2 Sources

    Teortaxes▶️ (DeepSeek 推特🐋铁粉 2023 – ∞)@teortaxesTexI will also be mildly surprised if we train 100T MoEs with dumb routers. It can clearly be done but it feels like such a waste1h
    will brown@willcbone of the most exciting directions here is the MCMC stuff imo, felt like the big missing idea in bridging midtraining -> RL but “credit assignment” can’t be generic without bias, needs to allow compute scaling, and can’t assume “partial progress” is well defined token-level is dead in the water, sampled policy gradient for verifiable tasks is sometimes just optimal though if “partial progress” is measurable by eg a pairwise judge trained online, you can decomp long horizon -> episodic, prune progress checkpoint sets to winners, branch again at each episode, backprop judge scores via survival prediction etc34m