Better AI-agent scores can mask worse results on skill-retrieval tasks, a post warns
The post describes a paper’s proposed evaluation method, RAE: when an agent retrieves a skill, rerun that exact task without skill access and compare the results.
TLDR
A user summarizing “Skill Following: Evaluating Actual Skill Use in Retrieval-Enabled LLM Agents” argues that comparing tasks where an agent retrieved skills with tasks where it did not can be misleading. If the agent retrieves skills on easier problems, retrieval can look helpful even when it does nothing or makes the answer worse. According to the summary, the paper’s proposed fix, RAE, compares performance on the same task with and without skill access whenever retrieval occurs.
Combined views
1K
1 Source, first seen 23d ago