MiMo-V2.6 update targets repetitive tool calls
Xiaomiβs MiMo developers say the models sometimes repeated identical or similar tool calls, using up context and stalling tasks, especially in MiMo Desktop, MiMo Code and OpenCode.
TLDR
Xiaomiβs MiMo developers say a training penalty applied only after more than 32 tool calls per turn, leaving shorter bouts of repetition unchecked. They say they trained a repetition-focused reinforcement-learning teacher in 12 steps on about 7,000 examples and merged it into MiMo-V2.6 via MOPD at roughly 4% of the cost of a full mixRL retrain. The team reports that repetition dropped while benchmarks held steady, and says the updated models are open-sourced.
Combined views
54.4K
2 Sources, first seen 13h ago