• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI
    Announcement

    SWE-sweep aims to help build agents that find and fix bugs without guidance or supervision

    A post introducing SWE-sweep says top models score below 5% on the benchmark and calls it one of its creators’ most challenging.

    OP
    JY
    SL
    12 Sources, ,

    TLDR

    SWE-sweep is intended to help build agents that can find and fix bugs without guidance or supervision, according to a post from one of its creators. The post says top models score below 5% on the benchmark and describes it as one of the creators’ most challenging yet.

    Combined views

    11.5K

    12 Sources, first seen 4h ago

    likes

    Combined views

    11.5K

    12 Sources, first seen 4h ago

    201 likes
    4h ago
    first seen 4h ago
    201
    16 comments
    32 saves
    31 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Featured Source
    16 comments
    32 saves
    31 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    #9

    Today's Rank

    #9

    12 Sources

    @OfirPressSWE-sweep will enable building agents that can find and fix bugs with no guidance or supervision. Top models get <5%, so this is an ability that labs haven't started climbing yet. Our benchmarks set north stars for future AI, and this is one of our most challenging ones ever.
    @its_sserrotLove this new benchmark, I occasionally ask models to try and find bugs in my work and it's not surprising to see how low they score with limited context provided.
    @jyangballinMost SWE benchmarks tell an agent exactly what to fix + where to look. Imagine if an agent could instead *proactively* surface and fix issues themselves (before a developer even runs into them!) We evaluate this capability at scale. Introducing SWE-sweep! Led by 🐐 @KLieret
    @sten_sootlaCheck out SWE-sweep, our new coding benchmark! SWE-bench gives the model an issue to fix. SWE-sweep goes a step further, towards full autonomy: the issue description is gone - the model is dropped into a codebase and simply tasked to proactively find and fix as many latent bugs as possible.
    @StellaLisyProactivity is super important in real human-agent interactions especially as model capabilities exceed human performance, excited to see how to build proactive agents to solve SWE Sweep!
    @magpie_rayhouNew benchmark out! SWE-sweep is a new genre of coding agent benchmarks where users just give the agent a very broad task to work on (e.g., bug hunting) without further specific directions. I would imagine similar benchmarks out soon to proactively optimize latency, efficiency or do refactoring for a codebase. As always, great work from @KLieret and great team!

    12 Sources

    @OfirPressSWE-sweep will enable building agents that can find and fix bugs with no guidance or supervision. Top models get <5%, so this is an ability that labs haven't started climbing yet. Our benchmarks set north stars for future AI, and this is one of our most challenging ones ever.
    @its_sserrotLove this new benchmark, I occasionally ask models to try and find bugs in my work and it's not surprising to see how low they score with limited context provided.
    @jyangballinMost SWE benchmarks tell an agent exactly what to fix + where to look. Imagine if an agent could instead *proactively* surface and fix issues themselves (before a developer even runs into them!) We evaluate this capability at scale. Introducing SWE-sweep! Led by 🐐 @KLieret
    @sten_sootlaCheck out SWE-sweep, our new coding benchmark! SWE-bench gives the model an issue to fix. SWE-sweep goes a step further, towards full autonomy: the issue description is gone - the model is dropped into a codebase and simply tasked to proactively find and fix as many latent bugs as possible.
    @StellaLisyProactivity is super important in real human-agent interactions especially as model capabilities exceed human performance, excited to see how to build proactive agents to solve SWE Sweep!
    @magpie_rayhouNew benchmark out! SWE-sweep is a new genre of coding agent benchmarks where users just give the agent a very broad task to work on (e.g., bug hunting) without further specific directions. I would imagine similar benchmarks out soon to proactively optimize latency, efficiency or do refactoring for a codebase. As always, great work from @KLieret and great team!