A proposed method for ranking base models for coding agents before post-training
A post about an NVIDIA paper says its three rankings closely matched post-trained SWE-bench Verified scores across ten model pairs.
TLDR
A post describing an NVIDIA paper says five of six base models solved zero SWE-bench Verified tasks. The proposed method starts with tasks a strong post-trained agent solved, identifies the first code edit that makes the tests pass, then checks whether a base model can produce or recognize that fix using the preceding context. The post says three resulting rankings closely matched post-trained scores across ten model pairs, offering a way to choose checkpoints before post-training.
Combined views
395
1 Source, first seen ago
