Announcement
Qwen3.8-27B-Medium reportedly improves 11.3 points after 15 GRPO steps using 500 tasks
AfterQuery says its internal post-training team used tasks from the company’s SWE agent dataset for the experiment.
TLDR
Using 500 tasks from its SWE agent dataset, AfterQuery says its internal post-training team improved Qwen3.8-27B-Medium by 11.3 points in 15 steps of GRPO. A separate post from someone at the company says its in-house research team shows labs their models’ losses and provides data to close those gaps.
Combined views
18.5K
4 Sources, first seen ago
