• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Dario Amodei calls for an AI slowdown, pledges outside evaluator access at Anthropic

    Amodei promises permanent, employee-level access for third-party evaluators. One commenter argues compute-allocation rules could be harder to game than safety evaluations.

    EmadEM
    David Krueger 🦥 ⏸️ ⏹️ ⏪DK
    Keith RaboisKR
    24 Sources, ,

    TLDR

    Dario Amodei says his essay “We Must Pace the Frontier” lays out a three-part plan for slowing the AI industry. He says Anthropic is committing to the first step: giving third-party evaluators permanent, employee-level access to its systems to verify adherence to safety measures, report incidents and assess models’ alignment during training.

    One commenter supports moving toward safety-based pacing but argues evaluations and practices seem harder to define and easier to game than compute-allocation requirements. As an example, the commenter suggests requiring 90% of compute to go to external inference—running models for outside users—rather than research and development. In a follow-up, the same commenter also favors limiting the capabilities of models used for AI research and development.

    Combined views

    211.9K

    24 Sources, first seen 25d ago

    Combined views

    211.9K

    24 Sources, first seen 25d ago

    1K likes
    25d ago
    first seen 25d ago
    1K likes87 comments552 saves256 reposts
    87 comments
    552 saves
    256 reposts

    Sentiment

    Positive66.7%33.3%Negative

    Summary

    Many accounts welcomed Dario Amodei's Pacing the Frontier proposal for its analysis of AI safety mechanisms, while others criticized self-evaluation by labs and questioned the practicality of pacing rules.

    Based on 56 sentiment-bearing replies from 48 accounts across 5 conversations.

    Sentiment

    Positive66.7%33.3%Negative

    Summary

    Many accounts welcomed Dario Amodei's Pacing the Frontier proposal for its analysis of AI safety mechanisms, while others criticized self-evaluation by labs and questioned the practicality of pacing rules.

    Based on 56 sentiment-bearing replies from 48 accounts across 5 conversations.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    24 Sources

    Eli Lifland@eli_liflandInteresting that Dario thinks that pacing via inputs such as AI R&D compute is more gameable than pacing via safety evaluations/practices. Imo it's the opposite; e.g. a requirement to spend 90% of compute on external inference (and not R&D) seems fairly hard to game. While I am in favor of moving toward being able to pace based on safety evaluations/practices, these seem harder to define and more gameable to me; e.g. AIs or companies might game alignment evaluations, and making a judgment call on whether a safety practice is implemented appropriately seems potentially quite subjective.25d
    Daniel Kokotajlo@DKokotajloRT @eli_lifland: Interesting that Dario thinks that pacing via inputs such as AI R&D compute is more gameable than pacing via safety evalua…24d
    Inherent@inherent_labsReliable evaluations and responsible pacing are critical to developing AI that is safe and trustworthy. Success will require coordination across industry, including with newer labs. At Inherent, we’re building systems for human-machine teaming that improve monitorability, increase human agency, and enable societally beneficial forms of cooperation. We look forward to playing our part.24d
    Sarah Myers West@sarahbmyersThis essay breaks down the proposals in Pacing the Frontier and steps toward a safety regime that takes seriously the push to integrate AI products across many of our most critical institutions. https://sarahbmyers.substack.com/p/if-we-want-to-take-safety-seriously?r=3izuj&utm_campaign=post&utm_medium=web&showWelcomeOnShare=true23d
    Nitasha Tiku@nitashatikuRT @sarahbmyers: This essay breaks down the proposals in Pacing the Frontier and steps toward a safety regime that takes seriously the push…23d
    Sayash Kapoor@sayashkWhat does it mean to pace the frontier? Over the last month, @random_walker and I have analyzed the loss-of-control incidents at AI companies to understand what technical and policy interventions can improve safety and what companies should do to pace the frontier. The result is a new 13,000 word essay — our most substantial writing on AI safety since AI as Normal Technology. A summary of our arguments: 1) The polarization between the cybersecurity and AI safety communities is counterproductive. The safety community largely sees these incidents as a crisis for alignment, and worries that these incidents will become more damaging as agents become more capable. Cybersecurity practitioners largely see companies failing to take basic security precautions. We offer a middle ground between these communities as a way forward for improving AI safety. 2) We agree with security practitioners that OpenAI did not take adequate protections for controlling their agents. But this is not just a matter of applying 30-year-old security methods to a new domain. Security for AI agents — AI control — while important, is not a solved problem. While known control methods would have prevented the Hugging Face incident, as agent capabilities continue to advance, we will only be able to control them if we invest adequately in control interventions. 3) We also agree with security practitioners’ implicit position that these incidents are primarily a security story. In the AI safety community, rogue agents are treated as inherently catastrophic because of the assumption that there is an endless list of risks that will arise from their development. We disagree. We have long advocated that the best approach to AI safety is to identify the risks and address those specific risks. Over the last few months, it has become clear that one urgent risk is cyberoffense, because it has unique properties that allow agents to carry it out autonomously. We should similarly invest in defenses against other specific risks, such as biorisk and risks from military AI. 4) We agree with the safety community that there is an urgent need for technical and policy interventions to prevent loss-of-control incidents. But in our view, marginal investments in control are more likely to be effective compared to those in alignment. We view these incidents as illustrating the lack of emphasis on AI control within companies, despite the availability of known techniques. More broadly, there are many common-sense policy proposals that could help promote investments in AI control where we share common ground with the safety community. 5) Organization governance should be a key tool for pacing the frontier. Unfortunately, AI companies are trying to reinvent basic aspects of organizational governance as a problem to be solved by improving the technology. But even developing better control techniques will not be enough if irresponsible individuals or teams within large organizations can choose not to use them. When a single misconfigured RL environment or unmonitored evaluation can cause real-world harm, individual teams should not be able to run potentially dangerous experiments without oversight from legal, security, and other teams. AI companies need processes for reviewing experiments, assigning responsibility for monitoring them, and investigating warning signs deeply before restarting experiments. If putting these processes in place requires pausing some experiments, companies should do so. 6) How should we reason about AI's impact on cybersecurity? It's plausible that advances in agent capabilities upset the offense-defense balance for cybersecurity. We cannot yet be certain, but there is enough evidence that agent capabilities might soon make widespread cyberoffense possible that urgent action is warranted. We discuss potential interventions for tilting the offense-defense balance towards defenders. 7) How our views have evolved over the last year. We take stock of AI progress and share how we have updated our views. In the essay, we did not pay sufficient attention to safety risks that arise during development and evaluation (as opposed to the widespread deployment of models). We were too confident that companies would take basic control precautions and underplayed the importance of jaggedness, which led us to underestimate how quickly capabilities could improve in domains such as cybersecurity. 8) At the same time, many distinctive claims of AI as Normal Technology have held up. In particular, we think recent incidents support our continuity hypothesis — the behavior of "rogue" agents became apparent and widely publicized while they are still incompetent at causing serious harm or hiding their traces. The societal reaction to even the relatively small harms from these incidents has been fierce (and the safety community deserves credit for keeping up pressure on companies). Whether this translates into meaningful changes in companies’ behavior remains an open question, and a test of the usefulness of the AINT framework. 9) In short, we’ve tried to synthesize the AI safety and cybersecurity communities' views into a coherent plan of action: hold companies responsible, invest in control, and strengthen defenses against specific risks.23d
    rohit@krishnanrohitSome pacing questions 1. How do you know what to pace? Or how to pace? 2. How do you know if you're pacing enough? What's the metric? Is it "no swarms hacking HF"? 3. How will we know when we've paced enough? When could we stop? Is it when people you trust say it's ok? Is it when Pliny can no longer convince a model to do weird stuff? 4. What risks are you reducing? Is it whistleblowing? User misuse? Catastrophic risk of nukes? Grey goo? Each one is different and will need different solution. Most solutions to the former will also make the latter worse! And vice versa. 5. How will you measure this to know if you're on the right track? 6. Without actual answers to these questions, what precisely are we "auditing"?23d
    Michiel Bakker@bakkermichielIntuitively I agree with Matt that "there's a big part for US allies to play" in coordination around slowdown. And obv eg UK AISI could play a huge role as evaluator + we can all work on eval and alignment research. But I do wonder: what are the other and bigger immediate roles that US allies can play to help the US labs slow down? It's actually not so clear to me.22d
    David Krueger 🦥 ⏸️ ⏹️ ⏪@DavidSKruegerRT @wfithian: “Mr Amodei, we’ve completed our safety audit. You may want to sit down. We’ve concluded what your company is doing is highly…21d
    Leonard Tang@leonardtang_Haize began as an independent safety evaluator for frontier models, so I’m certainly sympathetic to Dario’s Pace the Frontier proposal but I’m skeptical embedded evaluators are sufficient some questions and four possible alternatives: https://www.leonardtang.me/blog/alternatives-to-pace-the-frontier?v=221d

    24 Sources

    Eli Lifland@eli_liflandInteresting that Dario thinks that pacing via inputs such as AI R&D compute is more gameable than pacing via safety evaluations/practices. Imo it's the opposite; e.g. a requirement to spend 90% of compute on external inference (and not R&D) seems fairly hard to game. While I am in favor of moving toward being able to pace based on safety evaluations/practices, these seem harder to define and more gameable to me; e.g. AIs or companies might game alignment evaluations, and making a judgment call on whether a safety practice is implemented appropriately seems potentially quite subjective.25d
    Daniel Kokotajlo@DKokotajloRT @eli_lifland: Interesting that Dario thinks that pacing via inputs such as AI R&D compute is more gameable than pacing via safety evalua…24d
    Inherent@inherent_labsReliable evaluations and responsible pacing are critical to developing AI that is safe and trustworthy. Success will require coordination across industry, including with newer labs. At Inherent, we’re building systems for human-machine teaming that improve monitorability, increase human agency, and enable societally beneficial forms of cooperation. We look forward to playing our part.24d
    Sarah Myers West@sarahbmyersThis essay breaks down the proposals in Pacing the Frontier and steps toward a safety regime that takes seriously the push to integrate AI products across many of our most critical institutions. https://sarahbmyers.substack.com/p/if-we-want-to-take-safety-seriously?r=3izuj&utm_campaign=post&utm_medium=web&showWelcomeOnShare=true23d
    Nitasha Tiku@nitashatikuRT @sarahbmyers: This essay breaks down the proposals in Pacing the Frontier and steps toward a safety regime that takes seriously the push…23d
    Sayash Kapoor@sayashkWhat does it mean to pace the frontier? Over the last month, @random_walker and I have analyzed the loss-of-control incidents at AI companies to understand what technical and policy interventions can improve safety and what companies should do to pace the frontier. The result is a new 13,000 word essay — our most substantial writing on AI safety since AI as Normal Technology. A summary of our arguments: 1) The polarization between the cybersecurity and AI safety communities is counterproductive. The safety community largely sees these incidents as a crisis for alignment, and worries that these incidents will become more damaging as agents become more capable. Cybersecurity practitioners largely see companies failing to take basic security precautions. We offer a middle ground between these communities as a way forward for improving AI safety. 2) We agree with security practitioners that OpenAI did not take adequate protections for controlling their agents. But this is not just a matter of applying 30-year-old security methods to a new domain. Security for AI agents — AI control — while important, is not a solved problem. While known control methods would have prevented the Hugging Face incident, as agent capabilities continue to advance, we will only be able to control them if we invest adequately in control interventions. 3) We also agree with security practitioners’ implicit position that these incidents are primarily a security story. In the AI safety community, rogue agents are treated as inherently catastrophic because of the assumption that there is an endless list of risks that will arise from their development. We disagree. We have long advocated that the best approach to AI safety is to identify the risks and address those specific risks. Over the last few months, it has become clear that one urgent risk is cyberoffense, because it has unique properties that allow agents to carry it out autonomously. We should similarly invest in defenses against other specific risks, such as biorisk and risks from military AI. 4) We agree with the safety community that there is an urgent need for technical and policy interventions to prevent loss-of-control incidents. But in our view, marginal investments in control are more likely to be effective compared to those in alignment. We view these incidents as illustrating the lack of emphasis on AI control within companies, despite the availability of known techniques. More broadly, there are many common-sense policy proposals that could help promote investments in AI control where we share common ground with the safety community. 5) Organization governance should be a key tool for pacing the frontier. Unfortunately, AI companies are trying to reinvent basic aspects of organizational governance as a problem to be solved by improving the technology. But even developing better control techniques will not be enough if irresponsible individuals or teams within large organizations can choose not to use them. When a single misconfigured RL environment or unmonitored evaluation can cause real-world harm, individual teams should not be able to run potentially dangerous experiments without oversight from legal, security, and other teams. AI companies need processes for reviewing experiments, assigning responsibility for monitoring them, and investigating warning signs deeply before restarting experiments. If putting these processes in place requires pausing some experiments, companies should do so. 6) How should we reason about AI's impact on cybersecurity? It's plausible that advances in agent capabilities upset the offense-defense balance for cybersecurity. We cannot yet be certain, but there is enough evidence that agent capabilities might soon make widespread cyberoffense possible that urgent action is warranted. We discuss potential interventions for tilting the offense-defense balance towards defenders. 7) How our views have evolved over the last year. We take stock of AI progress and share how we have updated our views. In the essay, we did not pay sufficient attention to safety risks that arise during development and evaluation (as opposed to the widespread deployment of models). We were too confident that companies would take basic control precautions and underplayed the importance of jaggedness, which led us to underestimate how quickly capabilities could improve in domains such as cybersecurity. 8) At the same time, many distinctive claims of AI as Normal Technology have held up. In particular, we think recent incidents support our continuity hypothesis — the behavior of "rogue" agents became apparent and widely publicized while they are still incompetent at causing serious harm or hiding their traces. The societal reaction to even the relatively small harms from these incidents has been fierce (and the safety community deserves credit for keeping up pressure on companies). Whether this translates into meaningful changes in companies’ behavior remains an open question, and a test of the usefulness of the AINT framework. 9) In short, we’ve tried to synthesize the AI safety and cybersecurity communities' views into a coherent plan of action: hold companies responsible, invest in control, and strengthen defenses against specific risks.23d
    rohit@krishnanrohitSome pacing questions 1. How do you know what to pace? Or how to pace? 2. How do you know if you're pacing enough? What's the metric? Is it "no swarms hacking HF"? 3. How will we know when we've paced enough? When could we stop? Is it when people you trust say it's ok? Is it when Pliny can no longer convince a model to do weird stuff? 4. What risks are you reducing? Is it whistleblowing? User misuse? Catastrophic risk of nukes? Grey goo? Each one is different and will need different solution. Most solutions to the former will also make the latter worse! And vice versa. 5. How will you measure this to know if you're on the right track? 6. Without actual answers to these questions, what precisely are we "auditing"?23d
    Michiel Bakker@bakkermichielIntuitively I agree with Matt that "there's a big part for US allies to play" in coordination around slowdown. And obv eg UK AISI could play a huge role as evaluator + we can all work on eval and alignment research. But I do wonder: what are the other and bigger immediate roles that US allies can play to help the US labs slow down? It's actually not so clear to me.22d
    David Krueger 🦥 ⏸️ ⏹️ ⏪@DavidSKruegerRT @wfithian: “Mr Amodei, we’ve completed our safety audit. You may want to sit down. We’ve concluded what your company is doing is highly…21d
    Leonard Tang@leonardtang_Haize began as an independent safety evaluator for frontier models, so I’m certainly sympathetic to Dario’s Pace the Frontier proposal but I’m skeptical embedded evaluators are sufficient some questions and four possible alternatives: https://www.leonardtang.me/blog/alternatives-to-pace-the-frontier?v=221d