SkillLift reportedly reaches target performance with 40–70% fewer tokens
DAIR.AI describes a paper on SkillLift, which ranks edits to AI agents’ skill prompts with a learned rubric instead of requiring a full agent run to evaluate every revision.
TLDR
DAIR.AI says SkillLift learns to rank pairs of skill prompts using real agent outcomes. An inner loop revises prompts against that fixed rubric without new agent runs; an outer loop uses a few runs to recalibrate it. According to DAIR.AI’s summary, tests on SkillsBench and WildClawBench (147 tasks) with three models showed SkillLift beating SkillOpt and CoEvoSkills in all six combinations, even when those baselines had twice the token budget. The summary reports reaching target performance with 40–70% fewer tokens.
