• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI
    Reaction

    AI pretraining framed as selecting an action space and initial distribution

    In a reply, a user describes pretraining as “just selecting an action space and initial distribution.”

    Dimitris PapailiopoulosDP
    will brownWB
    Cameron R. Wolfe, Ph.D.CR
    3 Sources, ,

    TLDR

    A user looked back at something they said a year earlier. Another user replied with a terse take on AI pretraining: it’s “just selecting an action space and initial distribution.”

    Combined views

    2K

    3 Sources, first seen 4h ago

    likes

    Combined views

    2K

    3 Sources, first seen 4h ago

    47 likes
    4h ago
    first seen 4h ago
    47
    7 comments
    3 saves
    Featured Source
    7 comments
    3 saves

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    #14

    Today's Rank

    #14

    3 Sources

    will brown@willcb@yacineMTB pretraining is just selecting an action space and initial distribution4h
    Dimitris Papailiopoulos@DimitrisPapail@willcb @yacineMTB Pure RL seems so wasteful. I’d go as far as say RL is also v wasteful generically unless you’re at the absolute frontier and have no good supervision.2h
    Cameron R. Wolfe, Ph.D.@cwolferesearchexactly - we brute force on-policy trajectories for learning because collecting comparable human trajectories is too expensive / intractable. we generate way more rollouts for RL relative to supervised training. But, it's still less of a bottleneck to generate several OOMs more synthetic trajectories because we can scale it arbitrarily whereas human eval / annotation can't. so, supervised training basically gives the model a good enough prior to generate good enough synthetic data for continued training. there are also benefits to the fact that all rollouts are generated and, in turn, are more uniform / standardized compared to human trajectories1h

    3 Sources

    will brown@willcb@yacineMTB pretraining is just selecting an action space and initial distribution4h
    Dimitris Papailiopoulos@DimitrisPapail@willcb @yacineMTB Pure RL seems so wasteful. I’d go as far as say RL is also v wasteful generically unless you’re at the absolute frontier and have no good supervision.2h
    Cameron R. Wolfe, Ph.D.@cwolferesearchexactly - we brute force on-policy trajectories for learning because collecting comparable human trajectories is too expensive / intractable. we generate way more rollouts for RL relative to supervised training. But, it's still less of a bottleneck to generate several OOMs more synthetic trajectories because we can scale it arbitrarily whereas human eval / annotation can't. so, supervised training basically gives the model a good enough prior to generate good enough synthetic data for continued training. there are also benefits to the fact that all rollouts are generated and, in turn, are more uniform / standardized compared to human trajectories1h