AutoInteract reportedly improves coding and repository issue benchmarks
A thread reports gains on interactive LiveCodeBench-Pro and SWE-together, plus an improvement on standard LiveCodeBench-Pro for Qwen3.5-4B.
TLDR
The thread’s author reports large AutoInteract gains on an interactive LiveCodeBench-Pro variant with a user model and on SWE-together, a repository issue-resolution task involving a simulated user. They also report improved standard LiveCodeBench-Pro performance for Qwen3.5-4B. Their analysis says multi-turn interaction data, grounded user models and executable verification are necessary, and that reinforcement-learning training on AutoInteract tasks saturates less quickly and performs better than training on the source data.
Combined views
248
1 Source, first seen ago