Berkeley Paper: Outputs Alone Can Clone Agent Skills
Rohan Paul shares UC Berkeley research on reconstructing agent skills from outputs alone.
TLDR
Rohan Paul posted about a UC Berkeley paper titled Daydreaming. It describes an agent security issue where normal task outputs from an AI agent can expose enough behavioral data to let an attacker rebuild a portable copy of its hidden skill. The tweet states that these outputs recovered 86.8% of the behavior without needing the original prompt. Paul notes this creates a harder problem than prompt theft because the visible work itself leaks the capability.
Combined views
5.5K
2 Sources, first seen 28d ago