Tested image-based prompt attacks had at most 18% success with a realistic benign task present, team member says
A team member describes image-based prompt injection as less studied than text-based attacks. The evaluation covered existing methods targeting commercial vision-language models—AI systems that handle images and text.
TLDR
A team member reports that existing visual prompt injection methods they evaluated reached at most an 18% attack success rate against the commercial vision-language models they tested when a realistic, benign user task was present. They describe text-based prompt injection as well studied, but attacks through the image channel as less so.
Combined views
9.9K
2 Sources, first seen 23d ago