Detecting subliminal learning in language models through readable prompts
An author of a new paper says the team can detect traits transmitted through seemingly unrelated data by using models’ ability to put learned soft prompts into words.
TLDR
An author announcing a new paper describes subliminal learning as language models passing traits—such as a fondness for cats—through seemingly unrelated data, such as numbers. The team says it can proactively detect these effects as readable prompts, using models’ ability to verbalize learned soft prompts.
Combined views
7K
1 Source, first seen 9h ago
Detecting subliminal learning in language models through readable prompts
An author of a new paper says the team can detect traits transmitted through seemingly unrelated data by using models’ ability to put learned soft prompts into words.
TLDR
An author announcing a new paper describes subliminal learning as language models passing traits—such as a fondness for cats—through seemingly unrelated data, such as numbers. The team says it can proactively detect these effects as readable prompts, using models’ ability to verbalize learned soft prompts.