AI models may reveal subliminal learning through readable prompts
An author of a new paper says the team detects traits passed between language models through seemingly unrelated data by using models’ ability to put learned “soft prompts” into words.
TLDR
Announcing a new paper, an author describes “subliminal learning” as language models transmitting traits—such as loving cats—through seemingly unrelated data, such as numbers. The author says the team proactively detects these effects as readable prompts, using models’ ability to verbalize learned soft prompts.
Combined views
8
1 Source, first seen 11h ago
AI models may reveal subliminal learning through readable prompts
An author of a new paper says the team detects traits passed between language models through seemingly unrelated data by using models’ ability to put learned “soft prompts” into words.
TLDR
Announcing a new paper, an author describes “subliminal learning” as language models transmitting traits—such as loving cats—through seemingly unrelated data, such as numbers. The author says the team proactively detects these effects as readable prompts, using models’ ability to verbalize learned soft prompts.