• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI
    Announcement

    LADE is designed to flag harmful LLM prompts before generation begins

    The researchers say LADE uses first-token probabilities to detect harmful prompts without accessing a model’s internal hidden states.

    Mohit BansalMB
    Vaidehi Patil @ COLM2026VP
    Wonjun LeeWL
    4 Sources, ,

    TLDR

    Researchers introducing LADE in a NeurIPS 2026 paper say it uses a k-nearest-neighbors classifier to identify harmful prompts before an LLM generates text. They report 96.6% average accuracy at distinguishing harmful from benign queries across the reference–target model pairs they tested. They also say it performs well against several jailbreak attacks while keeping over-refusal low.

    Combined views

    1.3K

    4 Sources, first seen 2h ago

    Combined views

    1.3K

    4 Sources, first seen 2h ago

    38 likes
    2h ago
    first seen 2h ago
    38 likes
    2 comments
    3 saves
    32 reposts
    2 comments
    3 saves
    32 reposts
    Featured Source

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    4 Sources

    Vaidehi Patil @ COLM2026@vaidehi_patil_🎉 Excited to share our #NeurIPS2026 paper, "Safeguarding LLMs via Model-Agnostic Latent Safety Signals from Dark Knowledge"! Latent Safety Signals in LLMs: Can we detect a harmful query before the model generates even one token? We show that the answer lies in the model's dark knowledge, i.e., the first-token probability distribution contains safety signals that transfer across different LLMs. We use this insight to build LADE (Latent Safety Signals for Defense), a model-agnostic defense against jailbreak attacks that requires no access to internal hidden states. 🛡️ Key takeaways: 1️⃣ We discover latent safety signals in the first-token probability distribution: contrasting harmful and benign queries reveals safety-discriminative tokens that transfer consistently across safety-aligned LLMs, despite different architectures, tokenizers, and refusal styles. 2️⃣ We introduce LADE 🔍: a lightweight framework that identifies harmful queries before generation begins by applying a simple kNN classifier to latent safety signals in the first-token probability distribution. 3️⃣ LADE also transfers across models 🔄. Across all reference–target LLM pairs, it achieves 96.6% average harmful-vs-benign classification accuracy. It also performs well against several jailbreak attacks while keeping over-refusal low.2h
    Wonjun Lee@wonjun_lee_Excited to share our #NeurIPS2026 paper "Safeguarding LLMs via Model-Agnostic Latent Safety Signals from Dark Knowledge"! 🎉 We show that LLMs contain latent safety signals in their first-token probability distributions and these signals can transfer across different models.2h
    Mohit Bansal@mohitban47RT @vaidehi_patil_: 🎉 Excited to share our #NeurIPS2026 paper, "Safeguarding LLMs via Model-Agnostic Latent Safety Signals from Dark Knowle…2h

    4 Sources

    Vaidehi Patil @ COLM2026@vaidehi_patil_🎉 Excited to share our #NeurIPS2026 paper, "Safeguarding LLMs via Model-Agnostic Latent Safety Signals from Dark Knowledge"! Latent Safety Signals in LLMs: Can we detect a harmful query before the model generates even one token? We show that the answer lies in the model's dark knowledge, i.e., the first-token probability distribution contains safety signals that transfer across different LLMs. We use this insight to build LADE (Latent Safety Signals for Defense), a model-agnostic defense against jailbreak attacks that requires no access to internal hidden states. 🛡️ Key takeaways: 1️⃣ We discover latent safety signals in the first-token probability distribution: contrasting harmful and benign queries reveals safety-discriminative tokens that transfer consistently across safety-aligned LLMs, despite different architectures, tokenizers, and refusal styles. 2️⃣ We introduce LADE 🔍: a lightweight framework that identifies harmful queries before generation begins by applying a simple kNN classifier to latent safety signals in the first-token probability distribution. 3️⃣ LADE also transfers across models 🔄. Across all reference–target LLM pairs, it achieves 96.6% average harmful-vs-benign classification accuracy. It also performs well against several jailbreak attacks while keeping over-refusal low.2h
    Wonjun Lee@wonjun_lee_Excited to share our #NeurIPS2026 paper "Safeguarding LLMs via Model-Agnostic Latent Safety Signals from Dark Knowledge"! 🎉 We show that LLMs contain latent safety signals in their first-token probability distributions and these signals can transfer across different models.2h
    Mohit Bansal@mohitban47RT @vaidehi_patil_: 🎉 Excited to share our #NeurIPS2026 paper, "Safeguarding LLMs via Model-Agnostic Latent Safety Signals from Dark Knowle…2h