• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI
    Report

    All existing distillation-attack defenses are claimed to be evaluated after distillation

    The original author makes the claim while sharing recent work in a reposted message.

    YA
    SJ
    2 Sources, 13h ago, first seen 13h ago

    TLDR

    An author sharing recent work claims that all existing defenses against distillation attacks are evaluated after distillation. The message was reposted without added commentary.

    Combined views

    4.4K

    2 Sources, first seen 13h ago

    44 likes

    Combined views

    4.4K

    2 Sources, first seen 13h ago

    44 likes
    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    2 comments
    25 saves
    8 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Featured Source
    2 comments
    25 saves
    8 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    2 Sources

    @shidan_javaheriExcited to share our recent work! All existing defenses against distillation attacks are evaluated after distillation, implicitly assuming no further RL training. We find this makes existing evals give a false sense of security, and that RL makes simple attacks effective
    @yaringalRT @shidan_javaheri: Excited to share our recent work! All existing defenses against distillation attacks are evaluated after distillation…

    2 Sources

    @shidan_javaheriExcited to share our recent work! All existing defenses against distillation attacks are evaluated after distillation, implicitly assuming no further RL training. We find this makes existing evals give a false sense of security, and that RL makes simple attacks effective
    @yaringalRT @shidan_javaheri: Excited to share our recent work! All existing defenses against distillation attacks are evaluated after distillation…