• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI
    Report

    A SWE-2 reward-function extension aims to steer score-cost trade-offs

    The author says adjusting Cognition’s SWE-2 cost weight can target a desired balance between score and cost improvements.

    T(
    AL
    2 Sources, ,

    TLDR

    The author says they extended Cognition’s SWE-2 reward function to steer the Pareto frontier of score-cost trade-offs, rather than just improve it. SWE-2 uses S − λₑC, with λₑ set to the frontier’s derivative at a given effort. They say adapting λₑ can target a desired balance between score and cost improvements.

    Combined views

    41.2K

    2 Sources, first seen 1d ago

    Combined views

    41.2K

    2 Sources, first seen 1d ago

    331 likes
    1d ago
    first seen 1d ago
    331 likes
    22 comments
    254 saves
    52 reposts
    Featured Source
    22 comments
    254 saves
    52 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    2 Sources

    @_anishlkWe extend @cognition’s SWE-2 reward function to steer the Pareto frontier, not just improve it. SWE-2 uses S − λₑC, where λₑ is cleverly the frontier’s derivative at that effort. We show that adapting λₑ can target a desired score/cost improvement balance. See below!1d
    @teortaxesTexRT @_anishlk: We extend @cognition’s SWE-2 reward function to steer the Pareto frontier, not just improve it. SWE-2 uses S − λₑC, where λₑ…21h

    2 Sources

    @_anishlkWe extend @cognition’s SWE-2 reward function to steer the Pareto frontier, not just improve it. SWE-2 uses S − λₑC, where λₑ is cleverly the frontier’s derivative at that effort. We show that adapting λₑ can target a desired score/cost improvement balance. See below!1d
    @teortaxesTexRT @_anishlk: We extend @cognition’s SWE-2 reward function to steer the Pareto frontier, not just improve it. SWE-2 uses S − λₑC, where λₑ…21h