Report
MiMo-V2.6 reportedly uses AI agents to build tasks and grade answers during training
A user sharing the paper says humans set the budget and rules while agents audit tests and hunt for cheats.
TLDR
A user sharing the MiMo-V2.6 paper says AI agents build training tasks, audit tests, grade answers and look for cheats, while humans set the budget and rules. The user says a grader agent rewarded cleaner patches that passed tests rather than workarounds. MiMo-V2.6-Pro’s DeepSWE score reportedly rose from 58.4 to 72.6 over $2.6 million of reinforcement learning and was still climbing when training stopped.
Combined views
—
1 Source, first seen ago
— likes— comments— saves— reposts
