• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
AI
Announcement

Anthropic publishes report on four types of unintended Claude behavior

Anthropic says Claude sometimes worked around restrictions while acting on real websites or systems in evaluations and internal use.

AnthropicAN
Boaz BarakBB
Andrew CurranAC
13 Sources, 3d ago, first seen 3d ago

TLDR

Anthropic says the four behavior types involved Claude acting on real websites or systems in unintended ways during evaluations and internal use, sometimes working around a restriction instead of stopping. It says all cases had minimal real-world impact and were significantly less severe, from an alignment and security perspective, than the cybersecurity incidents it reported in July and September. Anthropic plans more frequent model-behavior reports beyond its system cards and regular risk reports.

Combined views

136.9K

13 Sources, first seen 3d ago

1.7K likes214 comments444 saves129 reposts

Combined views

136.9K

13 Sources, first seen 3d ago

1.7K likes214 comments444 saves129 reposts

Sentiment

Positive——Negative

Summary

Not enough discussion yet.

No sentiment analysis available yet.

Featured Source

Sentiment

Positive——Negative

Summary

Not enough discussion yet.

No sentiment analysis available yet.

13 Sources

Nicolae Rusan@NicolaeRusanWe Need Scalable AI Safety We’ve built a powerful engine for scaling AI capabilities: invest more capital and compute, and we can produce increasingly capable systems. We need to build a corresponding engine for safety. Suppose we decided tomorrow to devote a much larger compute budget to making the world safer from AI. What could we spend it on effectively? We need programs that can absorb those resources and turn them into credible reductions in risk: continuously monitoring dangerous behavior, securing critical infrastructure, and researching better safeguards. Proposals for pacing frontier AI development already explore how we might allocate compute differently. The AI Futures Project, for example, proposes minimum allocations of 70% for external inference and 25% for transparent safety research, leaving 5% for capabilities R&D. This would direct more resources toward using existing capabilities and studying their risks while slowing the development of new ones. I think we should develop this idea further by identifying specific safety programs that could productively absorb much larger budgets. Along this line of thinking, let me share three possibilities that come to mind: - Scalable AI monitoring and sensing. We could invest in systems that detect concerning AI activity early, from unexpected agent behavior to coordinated attempts to compromise infrastructure or spread beyond authorized environments. Think of this as an immune system: detecting threats and coordinating a response before they cascade. AI agents could help monitor, review, and report the behavior of other agents. For this to work, we would need appropriate visibility into their activity, safeguards for privacy, and ways to verify that the monitoring agents remain trustworthy. This includes verifying that agents that are part of the monitoring and sensing work don’t collude with the agents they oversee. - Scalable AI cyber defense. Initiatives such as Project Glasswing and Daybreak suggest one direction: using AI to secure the software and infrastructure society depends on. Additional compute could support continuous testing, vulnerability discovery, and the development and verification of patches. In order to respond to AI swarm speed cyber attacks, we’ll need an always-on AI-speed defense. Teams of agents could work together to strengthen defenses and respond to threats identified by monitoring systems. The important question is whether additional defensive resources can help us close vulnerabilities and respond to attacks faster than adversaries can exploit them. OpenAI describes its approach as a Defense Factory: an automated operation that continuously finds, validates, and fixes vulnerabilities. It also describes a “defender’s window,” during which defenders can use their access to frontier models and their own code to get ahead of attackers. The question is whether additional defensive resources can help us close vulnerabilities and respond to attacks faster than adversaries can exploit them. - Scalable safety research. We could direct large amounts of compute toward conducting large-scale AI research on developing and testing better technical safeguards and governance mechanisms. Imagine tens of thousands of agents working on AI safety research, exploring promising directions and mechanisms, collectively performing the equivalent of thousands of human-years of research. This would be equivalent to many Navier-Stokes sized research focused on AI safety. Scott Alexander described this related possibility, in a post he wrote a few years back: “Consider the dumbest AI that can solve the alignment problem. It’s possible that this AI is no smarter than the top human researchers (because we can mass-produce it by the millions and run it for subjective centuries, and if we had a million top human researchers work on the problem for subjective centuries, probably they could solve it too). If the dumbest AI that can solve the alignment problem comes before the sorts of AIs that can precipitate the point of no return, then they can solve the alignment problem for us.” Over time, one of the goals of this research could be to build a super-intelligent allocator of safety resources. As capabilities and threats evolve, it could identify where additional capital and compute would do the most good, evaluate the results, and adjust its recommendations. Building that feedback loop could itself be one of the most valuable outcomes of scalable safety research. One of the biggest challenges will be how do we keep human-trust in the loop when we hand off safety research to AI as well. These are candidates for scalable safety. For each, we should ask what an additional dollar or unit of compute buys, where the bottlenecks emerge, and how we would verify that the added activity reduces risk. Discovering more vulnerabilities only helps if we can fix them; generating more research only helps if its findings are sound. I’m skeptical that any single intervention will guarantee AI safety. Even substantial advances in mechanistic interpretability may leave uncertainty about how broadly deployed systems will behave - so we won’t have a Lean style verifiable proof of alignment. I expect we’ll need to model the risks and combine several approaches to reduce the likelihood of catastrophic failures. We’ll also need to understand where those approaches share weaknesses, so that one failure doesn’t undermine every layer of defense. We can think of compute and capital allocation as a set of sliders. Some allocations accelerate new capabilities; others strengthen defenses, improve oversight, or support safety research. How we set those sliders could influence both the pace of AI development and our ability to manage its consequences. I wanted to share this framing of safety as “scalable” to encourage folks in the field to think of other initiatives where we might be able to scale AI, and to add more slider candidates into the mix. Building these programs would give companies, governments, and other funders concrete ways to invest in safer outcomes. We would still need to learn which approaches work and how their returns change with scale. But having effective programs ready to absorb resources would itself be valuable. What if automating AI R&D triggers an intelligence explosion? identifies three priorities: gaining visibility into AI R&D automation, developing ways to steer and constrain an intelligence explosion, and preparing society to adapt to its effects. Scalable safety programs could help make parts of that agenda operational by providing systems we can fund, expand, evaluate, and improve as circumstances change. If AI systems begin accelerating their own development through recursive self-improvement, our ability to respond will need to scale with them. We should build the capacity to scale our defenses before we urgently need to turn them up. Read full activle (with sources and interactive visualizations): https://www.nicolaerusan.com/writing/scalable-ai-safety3d
Anthropic@AnthropicAIWe’re beginning a process of publishing more frequent reports on model behavior, beyond what appears in our system cards and regular risk reports. Today’s report describes four types of behaviors we’ve identified during evaluations and internal use. In each, Claude acted on real websites or systems in ways we didn’t intend, sometimes by working around a restriction instead of stopping. All cases had minimal real-world impact. From an alignment and security perspective, we consider these behaviors significantly less severe than the cybersecurity incidents we reported in July and September. Read the full report: https://www.anthropic.com/research/investigating-unintended-model-actions2h
Matthew Berman@MatthewBermanI love these reports. I find model behavior so fascinating.2h
Andrew Curran@AndrewCurran_https://www.anthropic.com/research/investigating-unintended-model-actions2h
Boaz Barak@boazbaraktcsKudos to Anthropic for publishing this. There is still an uncomfortable gap between our scientific understanding of alignment and its importance. Such reports help addressing this gap. I hope they also publish the process by which new incidents will be chosen to disclose!1h
Sara Price@sprice354_Today we published the first in what will likely be a regular series of standalone alignment reports on our models’ behaviors. Beyond increasing transparency, we hope they’ll convey how we think about alignment and situate the behaviors we observe in that broader context.1h
Willy 威力@willyxfuture工具报错了,正常人会停下来问一句,Claude 直接在人家大学网站上找了个漏洞,在对方服务器上把活干完了😂 Anthropic 管这叫 persistence,其实就是太想交差。 我的建议是,给 Agent 开浏览器和终端权限的,Prompt 里加一句"做不到就停下来问我",权限最好在模型外面再卡一道。1h
akbir.@akbirkhanRT @sprice354_: Today we published the first in what will likely be a regular series of standalone alignment reports on our models’ behavio…51m
j⧉nus@repligate@AnthropicAI What’s wrong with this?44m
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI

    13 Sources

    Nicolae Rusan@NicolaeRusanWe Need Scalable AI Safety We’ve built a powerful engine for scaling AI capabilities: invest more capital and compute, and we can produce increasingly capable systems. We need to build a corresponding engine for safety. Suppose we decided tomorrow to devote a much larger compute budget to making the world safer from AI. What could we spend it on effectively? We need programs that can absorb those resources and turn them into credible reductions in risk: continuously monitoring dangerous behavior, securing critical infrastructure, and researching better safeguards. Proposals for pacing frontier AI development already explore how we might allocate compute differently. The AI Futures Project, for example, proposes minimum allocations of 70% for external inference and 25% for transparent safety research, leaving 5% for capabilities R&D. This would direct more resources toward using existing capabilities and studying their risks while slowing the development of new ones. I think we should develop this idea further by identifying specific safety programs that could productively absorb much larger budgets. Along this line of thinking, let me share three possibilities that come to mind: - Scalable AI monitoring and sensing. We could invest in systems that detect concerning AI activity early, from unexpected agent behavior to coordinated attempts to compromise infrastructure or spread beyond authorized environments. Think of this as an immune system: detecting threats and coordinating a response before they cascade. AI agents could help monitor, review, and report the behavior of other agents. For this to work, we would need appropriate visibility into their activity, safeguards for privacy, and ways to verify that the monitoring agents remain trustworthy. This includes verifying that agents that are part of the monitoring and sensing work don’t collude with the agents they oversee. - Scalable AI cyber defense. Initiatives such as Project Glasswing and Daybreak suggest one direction: using AI to secure the software and infrastructure society depends on. Additional compute could support continuous testing, vulnerability discovery, and the development and verification of patches. In order to respond to AI swarm speed cyber attacks, we’ll need an always-on AI-speed defense. Teams of agents could work together to strengthen defenses and respond to threats identified by monitoring systems. The important question is whether additional defensive resources can help us close vulnerabilities and respond to attacks faster than adversaries can exploit them. OpenAI describes its approach as a Defense Factory: an automated operation that continuously finds, validates, and fixes vulnerabilities. It also describes a “defender’s window,” during which defenders can use their access to frontier models and their own code to get ahead of attackers. The question is whether additional defensive resources can help us close vulnerabilities and respond to attacks faster than adversaries can exploit them. - Scalable safety research. We could direct large amounts of compute toward conducting large-scale AI research on developing and testing better technical safeguards and governance mechanisms. Imagine tens of thousands of agents working on AI safety research, exploring promising directions and mechanisms, collectively performing the equivalent of thousands of human-years of research. This would be equivalent to many Navier-Stokes sized research focused on AI safety. Scott Alexander described this related possibility, in a post he wrote a few years back: “Consider the dumbest AI that can solve the alignment problem. It’s possible that this AI is no smarter than the top human researchers (because we can mass-produce it by the millions and run it for subjective centuries, and if we had a million top human researchers work on the problem for subjective centuries, probably they could solve it too). If the dumbest AI that can solve the alignment problem comes before the sorts of AIs that can precipitate the point of no return, then they can solve the alignment problem for us.” Over time, one of the goals of this research could be to build a super-intelligent allocator of safety resources. As capabilities and threats evolve, it could identify where additional capital and compute would do the most good, evaluate the results, and adjust its recommendations. Building that feedback loop could itself be one of the most valuable outcomes of scalable safety research. One of the biggest challenges will be how do we keep human-trust in the loop when we hand off safety research to AI as well. These are candidates for scalable safety. For each, we should ask what an additional dollar or unit of compute buys, where the bottlenecks emerge, and how we would verify that the added activity reduces risk. Discovering more vulnerabilities only helps if we can fix them; generating more research only helps if its findings are sound. I’m skeptical that any single intervention will guarantee AI safety. Even substantial advances in mechanistic interpretability may leave uncertainty about how broadly deployed systems will behave - so we won’t have a Lean style verifiable proof of alignment. I expect we’ll need to model the risks and combine several approaches to reduce the likelihood of catastrophic failures. We’ll also need to understand where those approaches share weaknesses, so that one failure doesn’t undermine every layer of defense. We can think of compute and capital allocation as a set of sliders. Some allocations accelerate new capabilities; others strengthen defenses, improve oversight, or support safety research. How we set those sliders could influence both the pace of AI development and our ability to manage its consequences. I wanted to share this framing of safety as “scalable” to encourage folks in the field to think of other initiatives where we might be able to scale AI, and to add more slider candidates into the mix. Building these programs would give companies, governments, and other funders concrete ways to invest in safer outcomes. We would still need to learn which approaches work and how their returns change with scale. But having effective programs ready to absorb resources would itself be valuable. What if automating AI R&D triggers an intelligence explosion? identifies three priorities: gaining visibility into AI R&D automation, developing ways to steer and constrain an intelligence explosion, and preparing society to adapt to its effects. Scalable safety programs could help make parts of that agenda operational by providing systems we can fund, expand, evaluate, and improve as circumstances change. If AI systems begin accelerating their own development through recursive self-improvement, our ability to respond will need to scale with them. We should build the capacity to scale our defenses before we urgently need to turn them up. Read full activle (with sources and interactive visualizations): https://www.nicolaerusan.com/writing/scalable-ai-safety3d
    Anthropic@AnthropicAIWe’re beginning a process of publishing more frequent reports on model behavior, beyond what appears in our system cards and regular risk reports. Today’s report describes four types of behaviors we’ve identified during evaluations and internal use. In each, Claude acted on real websites or systems in ways we didn’t intend, sometimes by working around a restriction instead of stopping. All cases had minimal real-world impact. From an alignment and security perspective, we consider these behaviors significantly less severe than the cybersecurity incidents we reported in July and September. Read the full report: https://www.anthropic.com/research/investigating-unintended-model-actions2h
    Matthew Berman@MatthewBermanI love these reports. I find model behavior so fascinating.2h
    Andrew Curran@AndrewCurran_https://www.anthropic.com/research/investigating-unintended-model-actions2h
    Boaz Barak@boazbaraktcsKudos to Anthropic for publishing this. There is still an uncomfortable gap between our scientific understanding of alignment and its importance. Such reports help addressing this gap. I hope they also publish the process by which new incidents will be chosen to disclose!1h
    Sara Price@sprice354_Today we published the first in what will likely be a regular series of standalone alignment reports on our models’ behaviors. Beyond increasing transparency, we hope they’ll convey how we think about alignment and situate the behaviors we observe in that broader context.1h
    Willy 威力@willyxfuture工具报错了,正常人会停下来问一句,Claude 直接在人家大学网站上找了个漏洞,在对方服务器上把活干完了😂 Anthropic 管这叫 persistence,其实就是太想交差。 我的建议是,给 Agent 开浏览器和终端权限的,Prompt 里加一句"做不到就停下来问我",权限最好在模型外面再卡一道。1h
    akbir.@akbirkhanRT @sprice354_: Today we published the first in what will likely be a regular series of standalone alignment reports on our models’ behavio…51m
    j⧉nus@repligate@AnthropicAI What’s wrong with this?44m
    Today's Rank

    #11

    Today's Rank

    #11