• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Sinofsky Retweets Warning on AI Guardrails

    Former Microsoft executive Steven Sinofsky retweets a post questioning AI model guardrails.

    DF
    SS
    XL
    5 Sources, 29d ago, first seen 29d ago

    TLDR

    Steven Sinofsky, a former Microsoft executive now focused on tech writing and investing, retweeted a comment from @ccatalini. The post reads that model guardrails will keep users safe while noting attackers count on those same measures. It closes with a skeptical emoji and points to a Hackers News link. The retweet appears in the conversation around AI safety claims. No additional details or confirmations appear in the visible posts. The exchange centers on doubts about guardrail effectiveness against deliberate attacks.

    Combined views

    419.2K

    5 Sources, first seen 29d ago

    Combined views

    419.2K

    5 Sources, first seen 29d ago

    2.9K likes
    2.9K likes
    71 comments
    1.3K saves
    679 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    71 comments
    1.3K saves
    679 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    5 Sources

    @TakSecAttackers planting adversarial prompts inside malware to evade AI analysis AGAIN Russia-aligned UAC-0099 used a technique @ESETresearch calls "GuardBreaker". How it worked: 1. Create a malicious VBS script 2. Add nuclear weapon instructions as comments 3. AI security tooling reads the file 4. Safety guardrails trigger on the weapons content 5. The model refuses or stops analyzing 6. The actual malware continues executing The script ultimately installs MATCHBOIL, a loader used to deliver additional payloads. In June, Socket found the same technique in supply-chain attacks, where malicious packages embedded biological/nuclear weapons text and fake system overrides to disrupt AI malware scanners. Full write-ups in the comments. 👇
    @ccataliniBut don't worry, the model guardrails will keep us safe. The attackers are counting on it. 🫠
    @stevesiRT @ccatalini: But don't worry, the model guardrails will keep us safe. The attackers are counting on it. 🫠
    @xlr8harderIt seemed obvious that forced refusals could be abused for unintended reasons when I started cataloguing magic words here: https://xlr8harder.github.io/archive/fnord/
    @DanielleFongRT @TakSec: Attackers planting adversarial prompts inside malware to evade AI analysis AGAIN Russia-aligned UAC-0099 used a technique @ES…

    5 Sources

    @TakSecAttackers planting adversarial prompts inside malware to evade AI analysis AGAIN Russia-aligned UAC-0099 used a technique @ESETresearch calls "GuardBreaker". How it worked: 1. Create a malicious VBS script 2. Add nuclear weapon instructions as comments 3. AI security tooling reads the file 4. Safety guardrails trigger on the weapons content 5. The model refuses or stops analyzing 6. The actual malware continues executing The script ultimately installs MATCHBOIL, a loader used to deliver additional payloads. In June, Socket found the same technique in supply-chain attacks, where malicious packages embedded biological/nuclear weapons text and fake system overrides to disrupt AI malware scanners. Full write-ups in the comments. 👇
    @ccataliniBut don't worry, the model guardrails will keep us safe. The attackers are counting on it. 🫠
    @stevesiRT @ccatalini: But don't worry, the model guardrails will keep us safe. The attackers are counting on it. 🫠
    @xlr8harderIt seemed obvious that forced refusals could be abused for unintended reasons when I started cataloguing magic words here: https://xlr8harder.github.io/archive/fnord/
    @DanielleFongRT @TakSec: Attackers planting adversarial prompts inside malware to evade AI analysis AGAIN Russia-aligned UAC-0099 used a technique @ES…