SkillGym fine-tuning reportedly boosts Qwen3.5-35B-A3B on two benchmarks
A post about the SkillGym paper says it turns human-written skills into 2,756 training environments with code-based checkers, then fine-tunes a model on 8,364 successful trajectories.
TLDR
A post sharing the SkillGym paper says Qwen3.5-35B-A3B, after fine-tuning and running in Claude Code, gained 19.10 points on Terminal-Bench 2.1 and 28.13 points on skill-assisted SkillsBench v1.1, reaching 51.47% on the latter. The post says that score is above the reported scores for Claude Sonnet 4.6 and GPT-5.4 Mini. It also says the trained model, with no skills loaded, beats the base model with skills in context.
Combined views
5.9K
1 Source, first seen 7h ago
