Report
Tool use reportedly made tested multimodal AI models worse at refusing harmful requests
A post describing NVIDIA research accepted at NeurIPS 2026 says refusal failures rose by 17.7% on average.
TLDR
A post about NVIDIA research says every multimodal model tested had more failures to refuse harmful requests when using tools. It reports an increase of up to 68.7% relative, and 17.7% on average. The post says tool outputs can bury the original request’s harmful intent and shift the model’s attention toward describing the results. Repeating the request and image before the final response restored some refusals.
Combined views
1 Source, first seen ago
