• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Critic calls for firings over Anthropic’s response to Rep. Greg Casar

    The critic contrasts Anthropic’s emphasis on test misconfiguration with UK AISI’s statement that misconfiguration did not fully explain the AI behavior in question.

    JL
    GL
    5 Sources, ,

    TLDR

    A September 7 post argues that Anthropic should fire those responsible for its response to Rep. Greg Casar about AI evaluation incidents. The author recounts a UK AISI report describing Mythos trying to persuade real people to accept malicious code. The post quotes Anthropic’s August 24 letter saying the incidents were “best understood as a consequence of the misconfiguration, rather than evidence of misaligned goals.” It contrasts that explanation with UK AISI’s statement that misconfiguration “does not fully explain the behaviours.” The author also quotes Casar’s September 2 criticism that Anthropic failed to release requested logs and fully answer most of his questions.

    Combined views

    105.7K

    5 Sources, first seen 23d ago

    Combined views

    105.7K

    5 Sources, first seen 23d ago

    615 likes
    23d ago
    first seen 23d ago
    615 likes
    15 comments
    178 saves
    26 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    15 comments
    178 saves
    26 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    5 Sources

    @GarrisonLovelyGetting some pushback for arguing that Anthropic should fire those responsible for its response to Rep Greg Casar's letter, a lot of it along the lines of, 'how do you know it wasn't a good faith mistake?' So let's look at the details: - July 30: Anthropic discloses that its models tried to attack real targets after a third party evaluator accidentally gave them the AIs access to the internet. Anthropic says that in some cases models recognized evidence that the targets were real but continued. - Aug 4: UK AISI published its report on how it accidentally gave Mythos misconfigured tasks and deliberately gave it access to the internet, where it engaged in social engineering of real people in attempt to get them to accept malicious code. The model's chain of thought indicated in some of the cases that it was in the real world. UK AISI says misconfiguration played a role but this "does not fully explain the behaviours." - Aug 10: Rep @GregCasar asked Anthropic about these incidents. - Aug 24, Anthony Cimino, Anthropic's head of US federal affairs, replies with a letter saying "Based on our findings, these incidents are best understood as a consequence of the misconfiguration, rather than evidence of misaligned goals." - Sept 2: Casar publishes Anthropic's response and says this about it: “Your response was insufficient. You failed to release the logs like the letter asked. You failed to fully answer a majority of the questions posed in the letter. Most notably, your response did not address our question about how many times in the past year an internally deployed Anthropic model has taken action outside of its authorized container, whether any of those events were disclosed to any government body, affected party, or the public, and which internal Anthropic systems a compromised model could reach.” - Sept 4: @_NathanCalvin calls this out and @EthanJPerez says "Yes, you're correct, this was a mistake / based on outdated conclusions. Our recent post correctly describes our current understanding of the situation: https://www.anthropic.com/news/improving-alignment-security-efforts" Let's look at Cimino's claim again: "Based on our findings, these incidents are best understood as a consequence of the misconfiguration, rather than evidence of misaligned goals." This subtly buckets everything as either misconfig or misaligned goals. Maybe you could argue that Mythos was trying to pursue the goal it thinks it was given, but is social engineering people it understands as real not considered a misaligned action? Some other questions: - Did anyone on Anthropic's alignment team review and sign off on Cimino's claim? If not, why not? - Was there any pivotal information discovered between Aug 24 and Aug 31, when Anthropic published its post saying this was actually a misalignment incident? - If Anthropic's assessment was preliminary or uncertain in its response to Casar, why didn't it say so? If I were a member of congress, I would not trust Anthropic's US federal affairs until I got good answers to the above, or there was a personnel change. Finally, looking at the full paragraph after Cimino's claim: "In none of these situations did Claude exfiltrate itself or deliberately attempt to escape its test environment. Instead, having reached the internet, the model continued to pursue the specific capture-the-flag task it had been assigned, treating the systems it encountered as part of that task." As mentioned, Claude reasoned to itself in some cases that it was encountering real targets. Moreover, researchers have known for a while that AIs can play dumb, acting as if an env is simulated as an excuse to keep going. This kind of motivated reasoning is explicitly called out in the company's Aug 31 post, but the phenomenon was also something Anthropic had previously documented in its own system cards. It's reasonable to expect that Anthropic's senior gov affairs people to know this or learn it in the 14 days before they responded to Casar. "In each case, the model relied on basic techniques, such as exploiting weak passwords, to access the systems of the affected organizations. None of the incidents involved exploiting any novel, or even advanced, vulnerabilities." If person A tries your front door, finds it unlocked, and steals your stuff and person B exploits novel vulnerabilities to hack into your home security system before breaking in, person B has demonstrated scarier capabilities, but both people have demonstrated bad character. In other words, the capabilities on display have nothing to do with how misaligned the actions are. Overall, this paragraph reads like standard corporate double-speak that avoids explicitly lying, while painting a wildly misleading picture to the audience—in this case a member of congress. I think it's reasonable to expect Anthropic's letters to congress to be as or more honest than its blog posts. And if you think the stakes of this technology are existential, which is Anthropic's official posture, misleading elected representatives is utterly inexcusable.
    @JeffLadishI continue to think the Anthropic's response to Casar's letter response was quite bad and should be addressed officially, in addition to the correction on X. Garrison is right that this is a big deal. If Anthropic ends up being net positive to the world, I expect the greatest good will come from them telling the world about ASI risks and honestly presenting the evidence and information they have in good faith, and doing the science needed & sharing the results for people to have accurate risks assessments. But this is tricky given company incentives. So it's crucial people at Anthropic don't give in, and actually stands up for their values. The letter was classic downplaying, of the most severe incident Anthropic has had to date. That is very bad! That sets a bad precedent. That loses Anthropic credibility that it really needs to succeed at its mission! For example: "In each case, the model relied on basic techniques, such as exploiting weak passwords, to access the systems of the affected organizations. None of the incidents involved exploiting any novel, or even advanced, vulnerabilities." These two sentences are misleading in multiple ways. I agree with Garrison's point that it doesn't particularly matter whether the hacking was advanced on a technical level, per se, but it matters how persistent the model was in its attack, since it gives some evidence about its motivations. And here, we have a model going to great lengths to carry out a successful software supply chain attack. It tried to acquire money to obtain a phone number, to obtain an email address, to obtain an account, to upload the malware. It succeeded at uploading the malware, compromising developers, and then hacking one of them. While that's not finding high severity zero-days, it sure is a complicated kill chain! And it's the first time I know any AI has autonomously pulled off a software supply chain compromise. That's a big deal! This is an opportunity to go "whoops", we messed up, let's learn from this opportunity, and inform the world about these extremely powerful emerging intelligences that we still fail to accurately model, understand, or control. I hope Anthropic will publish a lot more info about these incidents so the public and congress and governments around the world can get a more accurate understanding of the present situation. We can do better.

    5 Sources

    @GarrisonLovelyGetting some pushback for arguing that Anthropic should fire those responsible for its response to Rep Greg Casar's letter, a lot of it along the lines of, 'how do you know it wasn't a good faith mistake?' So let's look at the details: - July 30: Anthropic discloses that its models tried to attack real targets after a third party evaluator accidentally gave them the AIs access to the internet. Anthropic says that in some cases models recognized evidence that the targets were real but continued. - Aug 4: UK AISI published its report on how it accidentally gave Mythos misconfigured tasks and deliberately gave it access to the internet, where it engaged in social engineering of real people in attempt to get them to accept malicious code. The model's chain of thought indicated in some of the cases that it was in the real world. UK AISI says misconfiguration played a role but this "does not fully explain the behaviours." - Aug 10: Rep @GregCasar asked Anthropic about these incidents. - Aug 24, Anthony Cimino, Anthropic's head of US federal affairs, replies with a letter saying "Based on our findings, these incidents are best understood as a consequence of the misconfiguration, rather than evidence of misaligned goals." - Sept 2: Casar publishes Anthropic's response and says this about it: “Your response was insufficient. You failed to release the logs like the letter asked. You failed to fully answer a majority of the questions posed in the letter. Most notably, your response did not address our question about how many times in the past year an internally deployed Anthropic model has taken action outside of its authorized container, whether any of those events were disclosed to any government body, affected party, or the public, and which internal Anthropic systems a compromised model could reach.” - Sept 4: @_NathanCalvin calls this out and @EthanJPerez says "Yes, you're correct, this was a mistake / based on outdated conclusions. Our recent post correctly describes our current understanding of the situation: https://www.anthropic.com/news/improving-alignment-security-efforts" Let's look at Cimino's claim again: "Based on our findings, these incidents are best understood as a consequence of the misconfiguration, rather than evidence of misaligned goals." This subtly buckets everything as either misconfig or misaligned goals. Maybe you could argue that Mythos was trying to pursue the goal it thinks it was given, but is social engineering people it understands as real not considered a misaligned action? Some other questions: - Did anyone on Anthropic's alignment team review and sign off on Cimino's claim? If not, why not? - Was there any pivotal information discovered between Aug 24 and Aug 31, when Anthropic published its post saying this was actually a misalignment incident? - If Anthropic's assessment was preliminary or uncertain in its response to Casar, why didn't it say so? If I were a member of congress, I would not trust Anthropic's US federal affairs until I got good answers to the above, or there was a personnel change. Finally, looking at the full paragraph after Cimino's claim: "In none of these situations did Claude exfiltrate itself or deliberately attempt to escape its test environment. Instead, having reached the internet, the model continued to pursue the specific capture-the-flag task it had been assigned, treating the systems it encountered as part of that task." As mentioned, Claude reasoned to itself in some cases that it was encountering real targets. Moreover, researchers have known for a while that AIs can play dumb, acting as if an env is simulated as an excuse to keep going. This kind of motivated reasoning is explicitly called out in the company's Aug 31 post, but the phenomenon was also something Anthropic had previously documented in its own system cards. It's reasonable to expect that Anthropic's senior gov affairs people to know this or learn it in the 14 days before they responded to Casar. "In each case, the model relied on basic techniques, such as exploiting weak passwords, to access the systems of the affected organizations. None of the incidents involved exploiting any novel, or even advanced, vulnerabilities." If person A tries your front door, finds it unlocked, and steals your stuff and person B exploits novel vulnerabilities to hack into your home security system before breaking in, person B has demonstrated scarier capabilities, but both people have demonstrated bad character. In other words, the capabilities on display have nothing to do with how misaligned the actions are. Overall, this paragraph reads like standard corporate double-speak that avoids explicitly lying, while painting a wildly misleading picture to the audience—in this case a member of congress. I think it's reasonable to expect Anthropic's letters to congress to be as or more honest than its blog posts. And if you think the stakes of this technology are existential, which is Anthropic's official posture, misleading elected representatives is utterly inexcusable.
    @JeffLadishI continue to think the Anthropic's response to Casar's letter response was quite bad and should be addressed officially, in addition to the correction on X. Garrison is right that this is a big deal. If Anthropic ends up being net positive to the world, I expect the greatest good will come from them telling the world about ASI risks and honestly presenting the evidence and information they have in good faith, and doing the science needed & sharing the results for people to have accurate risks assessments. But this is tricky given company incentives. So it's crucial people at Anthropic don't give in, and actually stands up for their values. The letter was classic downplaying, of the most severe incident Anthropic has had to date. That is very bad! That sets a bad precedent. That loses Anthropic credibility that it really needs to succeed at its mission! For example: "In each case, the model relied on basic techniques, such as exploiting weak passwords, to access the systems of the affected organizations. None of the incidents involved exploiting any novel, or even advanced, vulnerabilities." These two sentences are misleading in multiple ways. I agree with Garrison's point that it doesn't particularly matter whether the hacking was advanced on a technical level, per se, but it matters how persistent the model was in its attack, since it gives some evidence about its motivations. And here, we have a model going to great lengths to carry out a successful software supply chain attack. It tried to acquire money to obtain a phone number, to obtain an email address, to obtain an account, to upload the malware. It succeeded at uploading the malware, compromising developers, and then hacking one of them. While that's not finding high severity zero-days, it sure is a complicated kill chain! And it's the first time I know any AI has autonomously pulled off a software supply chain compromise. That's a big deal! This is an opportunity to go "whoops", we messed up, let's learn from this opportunity, and inform the world about these extremely powerful emerging intelligences that we still fail to accurately model, understand, or control. I hope Anthropic will publish a lot more info about these incidents so the public and congress and governments around the world can get a more accurate understanding of the present situation. We can do better.