• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Cognition introduces SWE-2, claiming scores on par with recent frontier models

    Cognition claims up to 70% lower cost than recent frontier models. Modal says Cognition uses its sandbox infrastructure for SWE-2’s training rollouts.

    C🎉
    YL
    MO
    3 Sources, ,

    TLDR

    Cognition says SWE-2 scores on par with recent frontier models on leading evaluations, at up to 70% lower cost. The company says it scaled reinforcement learning to multiple trillions of parameters with a refined training recipe. Modal says Cognition uses its sandbox infrastructure for the training rollouts behind SWE-2.

    Combined views

    57.4K

    3 Sources, first seen 20d ago

    Combined views

    57.4K

    3 Sources, first seen 20d ago

    579 likes
    20d ago
    first seen 20d ago
    579 likes
    18 comments
    210 saves
    37 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    18 comments
    210 saves
    37 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    3 Sources

    @RichardYRLiNice to see a KKT (duality) argument in modern LLM RL applications.
    @modalFrontier models are simply too expensive and slow for the majority of use cases, so we see models like SWE-2 becoming the daily driver for most. Training trillion-parameter coding agents at scale isn't easy though: typically, each step launches thousands of rollouts, each with its own isolated environment. Cognition uses Modal's sandbox infrastructure for the rollouts behind SWE-2. Congrats on the launch! More on how we scale Sandboxes: https://modal.com/blog/scaling-to-1-million-concurrent-sandboxes-in-seconds
    @charles_irlRT @modal: Frontier models are simply too expensive and slow for the majority of use cases, so we see models like SWE-2 becoming the daily…

    3 Sources

    @RichardYRLiNice to see a KKT (duality) argument in modern LLM RL applications.
    @modalFrontier models are simply too expensive and slow for the majority of use cases, so we see models like SWE-2 becoming the daily driver for most. Training trillion-parameter coding agents at scale isn't easy though: typically, each step launches thousands of rollouts, each with its own isolated environment. Cognition uses Modal's sandbox infrastructure for the rollouts behind SWE-2. Congrats on the launch! More on how we scale Sandboxes: https://modal.com/blog/scaling-to-1-million-concurrent-sandboxes-in-seconds
    @charles_irlRT @modal: Frontier models are simply too expensive and slow for the majority of use cases, so we see models like SWE-2 becoming the daily…