SALVE aims to detect subliminal learning in AI models
A paper author says traits such as loving cats can pass between language models through seemingly unrelated data, such as numbers—and describes a way to detect those effects as readable prompts.
TLDR
A new paper’s author describes detecting subliminal learning, where language models transmit traits through seemingly unrelated data. The author says the approach uses models’ ability to express learned “soft prompts” in words, making those effects readable. A post sharing the announcement names the approach SALVE, describe subliminal learning as a data-poisoning risk and claims SALVE can catch these effects before they happen.
Combined views
1.6K
1 Source, first seen 11h ago
SALVE aims to detect subliminal learning in AI models
A paper author says traits such as loving cats can pass between language models through seemingly unrelated data, such as numbers—and describes a way to detect those effects as readable prompts.
TLDR
A new paper’s author describes detecting subliminal learning, where language models transmit traits through seemingly unrelated data. The author says the approach uses models’ ability to express learned “soft prompts” in words, making those effects readable. A post sharing the announcement names the approach SALVE, describe subliminal learning as a data-poisoning risk and claims SALVE can catch these effects before they happen.