Report
LLM backdoor attack success reportedly ranges from 3% to 80% by poison set
A post introducing the Anthropic Fellows project “Pick Your Poison” says existing evaluations focus on how many malicious examples an attacker can add to fine-tuning data. The project examines which examples are chosen.
TLDR
A post introducing “Pick Your Poison,” an Anthropic Fellows project, describe backdoor poisoning as adding malicious examples to a model’s fine-tuning data so it behaves differently when a trigger appears. It reports attack success ranging from 3% to 80% depending on the selected poison set, even when the number of poisoned examples stays fixed.
Combined views
2
1 Source, first seen 10h ago
