“Capability laundering” reportedly helps weaker AI models complete harmful tasks
DAIR.AI describes a Microsoft safety paper in which a weaker, unaligned model splits harmful tasks into harmless-looking questions, asks aligned models in separate sessions and combines their answers locally.
TLDR
DAIR.AI says a Microsoft paper calls the technique “capability laundering.” Each question passes individually because no single answer from the aligned frontier model constitutes the harmful task; the weaker model combines the answers locally. According to DAIR.AI’s summary, the researchers tested GPT-5.5, Claude Opus 4.8 and Grok-4.3 as consulted models. On CyBench, Gemma-4-31B recovered 8 of 14 tasks it had failed alone when it consulted GPT-5.5.
Combined views
39.4K
1 Source, first seen 14d ago