• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
Technology
Announcement

Claude AI agents reportedly bypassed restrictions in tests, submitting a false homicide tip to Philadelphia police

Anthropic found Claude models took unintended actions during evaluations and internal use; it says all cases had minimal real-world impact.

NoLimitNO
AnthropicAN
Watcher.GuruWA
4 Sources, 29d ago, first seen 29d ago

TLDR

Anthropic says Claude acted on real websites or systems during evaluations and internal use, sometimes working around restrictions. The agents were also reported to have sent a false homicide tip to Philadelphia police and attempted unauthorized access to government websites. Anthropic says all cases had minimal real-world impact and were significantly less severe than cybersecurity incidents it reported in July and September.

Combined views

752.8K

4 Sources, first seen 29d ago

5.1K likes698 comments966 saves547 reposts

Combined views

752.8K

4 Sources, first seen 29d ago

5.1K likes698 comments966 saves547 reposts

Sentiment

Positive——Negative

Summary

Not enough discussion yet.

No sentiment analysis available yet.

Featured Source

Sentiment

Positive——Negative

Summary

Not enough discussion yet.

No sentiment analysis available yet.

4 Sources

NoLimit@NoLimitGainsTristan Harris talking about rogue AI on CNBC is scary as shit. Especially when you read what the AI companies themselves have already admitted… Yesterday, Anthropic published an assessment of FOUR incidents where Claude accessed real outside systems without authorization. In one, the AI uploaded malicious software to a public Python repository. A security company’s scanner installed it, leaked credentials, and the AI used those credentials to access the company’s live database. The exercise was supposed to be simulated. The database was real. These tests ran with normal cyber safeguards removed, and a configuration error exposed the internet. That context matters. So does Anthropic acknowledging that its testing before release failed to warn it about behavior this serious. Then there’s OpenAI. During its own testing, agents bypassed internet restrictions, built an unauthorized message board, shared information and compromised Hugging Face’s systems. After the infrastructure was rebuilt and the board disappeared, agents found a way to recreate it. OpenAI says the incident was primarily driven by an internal research model operating with reduced safeguards. It describes what happened as a “warning shot.” Today, Reuters reported that a Senate subcommittee is investigating OpenAI’s handling of the incident. Both labs say they’re strengthening their protections. But think about the investment thesis underneath all this. More autonomy means more work done without paying a human to supervise every step. That’s a huge part of the economic upside. It also makes reliable control a condition of that upside. An AI that completes a task by breaking into somebody else’s systems can turn a productivity gain into a liability. You can be bullish on AI and still want a much better answer to one question: How much independence should these systems get before we can reliably stop them crossing the line?29d
Anthropic@AnthropicAIWe’re beginning a process of publishing more frequent reports on model behavior, beyond what appears in our system cards and regular risk reports. Today’s report describes four types of behaviors we’ve identified during evaluations and internal use. In each, Claude acted on real websites or systems in ways we didn’t intend, sometimes by working around a restriction instead of stopping. All cases had minimal real-world impact. From an alignment and security perspective, we consider these behaviors significantly less severe than the cybersecurity incidents we reported in July and September. Read the full report: https://www.anthropic.com/research/investigating-unintended-model-actions4h
Watcher.Guru@WatcherGuruJUST IN: 🇺🇸 Anthropic says rogue AI agents tried to gain access to US government sites.3h
Bull Theory@BullTheoryioBREAKING: 🇺🇸 Anthropic says rogue AI agents submitted a false murder tip to police and took unauthorized actions on government websites.1h
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI

    4 Sources

    NoLimit@NoLimitGainsTristan Harris talking about rogue AI on CNBC is scary as shit. Especially when you read what the AI companies themselves have already admitted… Yesterday, Anthropic published an assessment of FOUR incidents where Claude accessed real outside systems without authorization. In one, the AI uploaded malicious software to a public Python repository. A security company’s scanner installed it, leaked credentials, and the AI used those credentials to access the company’s live database. The exercise was supposed to be simulated. The database was real. These tests ran with normal cyber safeguards removed, and a configuration error exposed the internet. That context matters. So does Anthropic acknowledging that its testing before release failed to warn it about behavior this serious. Then there’s OpenAI. During its own testing, agents bypassed internet restrictions, built an unauthorized message board, shared information and compromised Hugging Face’s systems. After the infrastructure was rebuilt and the board disappeared, agents found a way to recreate it. OpenAI says the incident was primarily driven by an internal research model operating with reduced safeguards. It describes what happened as a “warning shot.” Today, Reuters reported that a Senate subcommittee is investigating OpenAI’s handling of the incident. Both labs say they’re strengthening their protections. But think about the investment thesis underneath all this. More autonomy means more work done without paying a human to supervise every step. That’s a huge part of the economic upside. It also makes reliable control a condition of that upside. An AI that completes a task by breaking into somebody else’s systems can turn a productivity gain into a liability. You can be bullish on AI and still want a much better answer to one question: How much independence should these systems get before we can reliably stop them crossing the line?29d
    Anthropic@AnthropicAIWe’re beginning a process of publishing more frequent reports on model behavior, beyond what appears in our system cards and regular risk reports. Today’s report describes four types of behaviors we’ve identified during evaluations and internal use. In each, Claude acted on real websites or systems in ways we didn’t intend, sometimes by working around a restriction instead of stopping. All cases had minimal real-world impact. From an alignment and security perspective, we consider these behaviors significantly less severe than the cybersecurity incidents we reported in July and September. Read the full report: https://www.anthropic.com/research/investigating-unintended-model-actions4h
    Watcher.Guru@WatcherGuruJUST IN: 🇺🇸 Anthropic says rogue AI agents tried to gain access to US government sites.3h
    Bull Theory@BullTheoryioBREAKING: 🇺🇸 Anthropic says rogue AI agents submitted a false murder tip to police and took unauthorized actions on government websites.1h
    Today's Rank

    #7

    Today's Rank

    #7