Paper Flags Bias in LLM Agent Skill Tests
Analysis of retrieval-enabled LLM agents shows standard tests mix unrelated tasks.
TLDR
DAIR.AI posted about work by Seonghyeon Cho and Chanjun Park at Korea University. The authors argue that common checks for skill retrieval in LLM agents compare retrieved-skill tasks against no-skill tasks. Because those tasks differ, the comparison mixes selection effects with any real skill benefit. They note that a skill improving an aggregate score can still reduce performance on every individual task it touches. The post links an arXiv paper titled Skill Following: Evaluating Actual Skill Use in Retrieval-Enabled LLM Agents.
Combined views
5.7K
1 Source, first seen 28d ago