Proposed multi-agent training targets repetitive LLM answers
The proposal’s authors say a single model takes on different roles that learn to give more varied answers—but only when those answers meet a quality threshold.
TLDR
The proposal’s authors say LLMs give strikingly similar answers to the same prompt, even across different model families. They propose multi-agent reinforcement learning during post-training, with different roles within one model learning to diversify responses only if those responses meet a quality threshold. The authors report improved quality and diversity over several competing diversity-enhancing methods across four benchmarks spanning scientific ideation and creative writing.
Combined views
13.7K
2 Sources, first seen 16d ago