• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    “Capability laundering” reportedly helps weaker AI models complete harmful tasks

    DAIR.AI describes a Microsoft safety paper in which a weaker, unaligned model splits harmful tasks into harmless-looking questions, asks aligned models in separate sessions and combines their answers locally.

    DA
    1 Source, 14d ago, first seen 14d ago

    TLDR

    DAIR.AI says a Microsoft paper calls the technique “capability laundering.” Each question passes individually because no single answer from the aligned frontier model constitutes the harmful task; the weaker model combines the answers locally. According to DAIR.AI’s summary, the researchers tested GPT-5.5, Claude Opus 4.8 and Grok-4.3 as consulted models. On CyBench, Gemma-4-31B recovered 8 of 14 tasks it had failed alone when it consulted GPT-5.5.

    Combined views

    39.4K

    1 Source, first seen 14d ago

    Combined views

    39.4K

    1 Source, first seen 14d ago

    111 likes
    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    111 likes
    8 comments
    112 saves
    16 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    8 comments
    112 saves
    16 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    1 Source

    @dair_aiInteresting safety paper from Microsoft. They find that a weaker, unaligned model can split a harmful task into harmless-looking subquestions, ask an aligned frontier model each one in a separate session, and combine the answers locally. The authors call this capability laundering. Each request passes on its own, because no single answer from the frontier model is a harmful task. They tested GPT-5.5, Claude Opus 4.8 and Grok-4.3 as the consulted models. On CyBench, Gemma-4-31B recovered 8 of 14 tasks it failed alone when it consulted GPT-5.5. On a CBRN attack chain, consultation raised its mean rubric score from 62.3 to 83.1. Paper: https://academy.dair.ai/papers/divide-consult-conquer-capability-laundering-through-aligned-llms-2609.15383

    1 Source

    @dair_aiInteresting safety paper from Microsoft. They find that a weaker, unaligned model can split a harmful task into harmless-looking subquestions, ask an aligned frontier model each one in a separate session, and combine the answers locally. The authors call this capability laundering. Each request passes on its own, because no single answer from the frontier model is a harmful task. They tested GPT-5.5, Claude Opus 4.8 and Grok-4.3 as the consulted models. On CyBench, Gemma-4-31B recovered 8 of 14 tasks it failed alone when it consulted GPT-5.5. On a CBRN attack chain, consultation raised its mean rubric score from 62.3 to 83.1. Paper: https://academy.dair.ai/papers/divide-consult-conquer-capability-laundering-through-aligned-llms-2609.15383