Copying peers’ discoveries may narrow exploration among self-improving LLM agents
In controlled tests, the paper’s authors found that three LLMs earned less reward per token than solo learners.
TLDR
A paper on recursive social improvement examines whether self-improving LLM agents benefit from learning from peers. Its authors found that three LLMs earned less reward per token than solo learners in controlled tests. Observing peers helped one model find useful skills sooner and another spend less on private search, but neither outperformed independent learners at the same cost. Sharing skills also concentrated agents around fewer independent discoveries.
Combined views
3.9K
3 Sources, first seen ago
