Kimi k2.7 improved on coding benchmarks after post-training, a user reports
The user says training on 1,700 Surge coding tasks made execution less brittle and reduced steps from 150 to 98 on DeepSWE and 102 to 78 on Terminal-Bench 3.
TLDR
A user involved in post-training Kimi k2.7 says the model had been dropping requirements, writing narrow tests and creating regressions. After reinforcement learning on 1,700 Surge coding tasks, they report benchmark gains of +20.0 on SWE-Marathon, +14.6 on Terminal-Bench 2.1, +12.4 on DeepSWE, +10.7 on Terminal-Bench 3 and +4.7 on SWE-Bench Pro. They also report fewer steps on DeepSWE and Terminal-Bench 3. According to the user, the training tasks weren't designed around these evaluations; DeepSWE, SWE-Marathon and Terminal-Bench 3 didn't exist when the tasks were collected.
Combined views
372
1 Source, first seen 19d ago