• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
AI
Reaction

At COLM, “science again” meets skepticism over RL loss tweaks

A post describes “science again” as a common refrain at COLM, while questioning whether RL loss tweaks are better.

(((ل()(ل() 'yoav))))👾('
Tal LinzenTL
5 Sources, 2h ago, first seen 2h ago

TLDR

A poster says a pervasive comment at COLM was that last year’s work was “boring” prompt tweaking and that researchers are now “doing science again.” The poster disputes that contrast, calling changes to reinforcement-learning loss terms or similar work “meaningless tweaks” and questioning why they are any better.

Combined views

6K

5 Sources, first seen 2h ago

88 likes10 comments19 saves6 reposts

Combined views

6K

5 Sources, first seen 2h ago

88 likes10 comments19 saves6 reposts

Sentiment

Positive——Negative

Summary

Not enough discussion yet.

No sentiment analysis available yet.

Featured Source

Sentiment

Positive——Negative

Summary

Not enough discussion yet.

No sentiment analysis available yet.

5 Sources

(((ل()(ل() 'yoav))))👾@yoavgoa pervasive statement at COLM was "last year was boring, we only did prompt tweaking, now we are doing science again". but said science is various meaningless tweaks to an RL loss term or similar, and I seriously don't see why its any better2h
Tal Linzen@tallinzenMismatch between people's training and identity as CS/ML researchers and what 90% of impactful LLM work actually is (data, evals, policy, applications).1h
Andrew Drozdov@mrdrozdovOne of my favorite directions is to match quality of on-policy RL with off-policy methods. Take a look at KARL and OAPL. More than tweaking :)29m
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI

    5 Sources

    (((ل()(ل() 'yoav))))👾@yoavgoa pervasive statement at COLM was "last year was boring, we only did prompt tweaking, now we are doing science again". but said science is various meaningless tweaks to an RL loss term or similar, and I seriously don't see why its any better2h
    Tal Linzen@tallinzenMismatch between people's training and identity as CS/ML researchers and what 90% of impactful LLM work actually is (data, evals, policy, applications).1h
    Andrew Drozdov@mrdrozdovOne of my favorite directions is to match quality of on-policy RL with off-policy methods. Take a look at KARL and OAPL. More than tweaking :)29m
    Today's Rank

    #13

    Today's Rank

    #13