Training Qwen on full Gemini runs reportedly cut success across seven tasks
A post summarizing Salesforce AI research says having Gemini correct Qwen’s own failed steps worked better than copying Gemini’s full runs: 79.7% success versus 63.1%.
TLDR
According to a post describing Salesforce AI research, tuning prompts, tools and workflow for Qwen3-Coder-30B-A3B raised average success from 29.2% to 78.0% across seven enterprise tasks. Gemini reached 93.6% in the same setup, but fine-tuning Qwen on complete Gemini runs cut its success to 63.1%, with declines on all seven tasks. The post attributes the drop to Qwen copying planning behavior that no longer fit its tuned setup. It says keeping Qwen’s own failed runs and having Gemini correct only the step where Qwen went wrong brought success to 79.7%.
