Reaction
At COLM, “science again” meets skepticism over RL loss tweaks
A post describes “science again” as a common refrain at COLM, while questioning whether RL loss tweaks are better.
TLDR
A poster says a pervasive comment at COLM was that last year’s work was “boring” prompt tweaking and that researchers are now “doing science again.” The poster disputes that contrast, calling changes to reinforcement-learning loss terms or similar work “meaningless tweaks” and questioning why they are any better.
Combined views
6K
5 Sources, first seen ago
88 likes10 comments19 saves6 reposts