Report
A SWE-2 reward-function extension aims to steer score-cost trade-offs
The author says adjusting Cognition’s SWE-2 cost weight can target a desired balance between score and cost improvements.
TLDR
The author says they extended Cognition’s SWE-2 reward function to steer the Pareto frontier of score-cost trade-offs, rather than just improve it. SWE-2 uses S − λₑC, with λₑ set to the frontier’s derivative at a given effort. They say adapting λₑ can target a desired balance between score and cost improvements.
Combined views
41.2K
2 Sources, first seen ago
