• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI
    Reaction

    The role of human oversight in automated pretraining research

    A user argues that close human supervision helps steer automated pretraining experiments toward production constraints.

    Minh Nhat Nguyen 🦭MN
    1 Source, 1h ago, first seen 1h ago

    TLDR

    A user says automated pretraining research yields more when people steer it toward a model’s production needs. They offer a hypothetical switch from softmax to sigmoid: even if loss falls faster, its effects on long context, post-training and other changes still need checking. They stress that autonomous pretraining research is theoretically solvable.

    Combined views

    1.3K

    1 Source, first seen 1h ago

    Combined views

    1.3K

    1 Source, first seen 1h ago

    20 likes
    20 likes
    3 comments
    13 saves
    1 reposts
    3 comments
    13 saves
    1 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    1 Source

    Minh Nhat Nguyen 🦭@menhguinmy experience w autoresearch for pretraining is that they 100% benefit from human-in-the-loop. you get way more out of them the closer you supervise and steer them towards constraints and desired direction. ultimately, pretraining research is meant to put a model into production, and you still need to verify and shape the model to these constraints. for example, let's say you implement sigmoid instead of softmax. okay, your loss went down faster. how does that changes your logits? does it fw long context? what happens if you stack it with 6 other things discovered in the autoresearch loop? does it have some weird ass impact on post-training or change based on ur pretraining data or ofc the scale of models. how do u troubleshoot all this new stuff you just added? this is all combinatorial complexity that if made much more efficient w human judgement. and additionally, you have to consider why you're being hired to make this model. what if it's a priority to be easier to serve and the you go back and forth for several days w the model and it isn't quite grokking what ur priorities are here, so it mode collapses in the ideas it suggests and ur back at square one by next standup. to be clear, this is not saying pretraining autoresearch can't be solved. it's a very theoretically solvable problem, and even if you spam infinite random walk ablations you can just pick 1 out of 100 that seems net-helpful.1h
    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    1 Source

    Minh Nhat Nguyen 🦭@menhguinmy experience w autoresearch for pretraining is that they 100% benefit from human-in-the-loop. you get way more out of them the closer you supervise and steer them towards constraints and desired direction. ultimately, pretraining research is meant to put a model into production, and you still need to verify and shape the model to these constraints. for example, let's say you implement sigmoid instead of softmax. okay, your loss went down faster. how does that changes your logits? does it fw long context? what happens if you stack it with 6 other things discovered in the autoresearch loop? does it have some weird ass impact on post-training or change based on ur pretraining data or ofc the scale of models. how do u troubleshoot all this new stuff you just added? this is all combinatorial complexity that if made much more efficient w human judgement. and additionally, you have to consider why you're being hired to make this model. what if it's a priority to be easier to serve and the you go back and forth for several days w the model and it isn't quite grokking what ur priorities are here, so it mode collapses in the ideas it suggests and ur back at square one by next standup. to be clear, this is not saying pretraining autoresearch can't be solved. it's a very theoretically solvable problem, and even if you spam infinite random walk ablations you can just pick 1 out of 100 that seems net-helpful.1h