Kimi K2.7 saw benchmark gains after training on 1,700 coding tasks, a user reports
The user describes the change as a move toward “shippable” code, saying the model previously dropped requirements, wrote narrow tests and created regressions.
TLDR
A user reports post-training Kimi K2.7 with reinforcement learning on 1,700 Surge coding tasks. They list benchmark gains of +20.0 on SWE-Marathon, +14.6 on Terminal-Bench 2.1, +12.4 on DeepSWE, +10.7 on Terminal-Bench 3 and +4.7 on SWE-Bench Pro.
Combined views
3.2K
1 Source, first seen 19d ago
likes