MiMo-V2.6 reportedly scales reinforcement learning to about 2 billion tokens per step
In a September 16 update, the team behind MiMo-V2.6 said it was scaling compute, agent environments and grading, with plans to open-source details in stages.
TLDR
The MiMo-V2.6 team said on September 16 that its reinforcement learning run was underway, using roughly 2 billion tokens per step in a fully asynchronous setup with 1,568 prompts × 16 rollouts. It said the run combines multiple agent tasks and harnesses, with rewards based on test cases and rubrics. The team shared a link to stream the run and said it would open-source details piece by piece over the following weeks.
Combined views
3.6M
26 Sources, first seen 1d ago
MiMo-V2.6 reportedly scales reinforcement learning to about 2 billion tokens per step
In a September 16 update, the team behind MiMo-V2.6 said it was scaling compute, agent environments and grading, with plans to open-source details in stages.