Agents trained against MIMESIS reportedly outperform those trained against GPT-5.5 under nine unseen user simulators
A post says Meta Superintelligence Labs trained its 9B simulator on human conversations and 13 real-user behavior patterns.
TLDR
A post describing a Meta Superintelligence Labs paper says agent training often relies on simulated users that are too cooperative and explicit. It says MIMESIS, a 9B model trained on human conversations and 13 real-user behavior patterns, beat Claude Opus 5 by 13.4 points on behavioral fidelity. Agents trained against MIMESIS outperformed those trained against GPT-5.5 under all nine unseen user simulators; turning its private reasoning into coaching feedback reportedly brought further gains.
Combined views
358
1 Source, first seen ago
