• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Kimi k2.7 improved on coding benchmarks after post-training, a user reports

    The user says training on 1,700 Surge coding tasks made execution less brittle and reduced steps from 150 to 98 on DeepSWE and 102 to 78 on Terminal-Bench 3.

    EC
    1 Source, 19d ago, first seen 19d ago

    TLDR

    A user involved in post-training Kimi k2.7 says the model had been dropping requirements, writing narrow tests and creating regressions. After reinforcement learning on 1,700 Surge coding tasks, they report benchmark gains of +20.0 on SWE-Marathon, +14.6 on Terminal-Bench 2.1, +12.4 on DeepSWE, +10.7 on Terminal-Bench 3 and +4.7 on SWE-Bench Pro. They also report fewer steps on DeepSWE and Terminal-Bench 3. According to the user, the training tasks weren't designed around these evaluations; DeepSWE, SWE-Marathon and Terminal-Bench 3 didn't exist when the tasks were collected.

    Combined views

    372

    1 Source, first seen 19d ago

    Combined views

    372

    1 Source, first seen 19d ago

    4 likes
    4 likes
    3 saves

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    3 saves

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    1 Source

    @echenthe data wasn’t designed around these evals. DeepSWE, SWE-Marathon, and Terminal-Bench 3 didn’t even exist when the tasks were collected. the trained model also used fewer steps, not more: 150 → 98 on DeepSWE, 102→78 on TB3. kimi k2.7 already knew how to write code, but post-training made its actual execution less brittle. Full writeup: https://surgehq.ai/blog/hill-climbing-swe-agent-kimi-k-2-7

    1 Source

    @echenthe data wasn’t designed around these evals. DeepSWE, SWE-Marathon, and Terminal-Bench 3 didn’t even exist when the tasks were collected. the trained model also used fewer steps, not more: 150 → 98 on DeepSWE, 102→78 on TB3. kimi k2.7 already knew how to write code, but post-training made its actual execution less brittle. Full writeup: https://surgehq.ai/blog/hill-climbing-swe-agent-kimi-k-2-7