Announcement
Qwen3.5-27B reportedly gains on unseen agent benchmarks after one training epoch
Toloka says the model’s τ³ retail score rose from 77.6% to 86.8% after training on its enterprise agent data.
TLDR
Toloka says it trained Qwen3.5-27B for one epoch on its enterprise agent data, then tested it on agent benchmarks the model had not seen. It reports a τ³ retail increase from 77.6% to 86.8%, AutomationBench gains of 3.4 percentage points and 9.6 points on Ops, and a 4.3-point Toolathlon gain. Toloka says the model learned to check policy before making changes.
Combined views
873
2 Sources, first seen ago
