• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Anthropic safety team member announces departure, urges more AI transparency

    The former team member says they plan to join METR to evaluate AI risks independently, arguing that AI companies are underinvesting in safety.

    GM
    TK
    SÓ
    6 Sources, ,

    TLDR

    On September 11, a former member of Anthropic's safety team said they had left two weeks earlier and planned to join METR for independent AI risk evaluations. They warned that a company could lose control of its systems without the public knowing. Their proposed guardrails include disclosing progress toward recursive self-improvement—AI systems improving themselves—reporting safety incidents and near-misses, and meeting minimum safety standards with independent guarantees that those standards are met.

    Combined views

    2.1M

    6 Sources, first seen 19d ago

    Combined views

    2.1M

    6 Sources, first seen 19d ago

    24K likes
    19d ago
    first seen 19d ago
    24K likes
    1.3K comments
    6.4K saves
    5.3K reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    1.3K comments
    6.4K saves
    5.3K reposts
    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    6 Sources

    @JoeJBentonI left Anthropic's safety team two weeks ago. Now feels like a good moment to explain why. AI companies are racing to build machines that are much smarter than any human, and we may not survive this. I want to work from the outside to ensure the public is informed about these risks, and help the world navigate this transition responsibly. Right now, AI companies are underinvesting in safety. A company could undergo an intelligence explosion, or lose control of its systems, without the public ever knowing. We only found out about the HuggingFace incident because the agents broke out onto the public internet. I don’t think that’s acceptable for a technology that might cause extinction-level risks. The public should demand far more transparency. We can’t steer this technology safely without more people being able to see where it’s going. Some of this is basic: companies should disclose their progress towards recursive self-improvement, report safety incidents and near-misses, meet minimum safety standards, and get independent guarantees that they are meeting those standards. I’ll be joining @METR_Evals to do independent evaluations of these risks. I want to show the world that these guardrails are possible, and that by doing them we can move these companies’ incentives away from racing and towards responsible development. I wrote up more thoughts here on my decision and what I hope changes: https://substack.com/@jbenton1/p-215263538
    @mattparlmerVery heartening to see calls for disclosure and transparency from former Anthropic staff, pretty clear that the era of heavily NDA’d research silos has caused immense damage to AI safety, regardless of your views on what the near term risks actually are
    @GaryMarcus“right now, AI companies are underinvesting in safety” might prove to be the understatement of the century.
    @S_OhEigeartaigh"Right now, AI companies are underinvesting in safety. A company could undergo an intelligence explosion, or lose control of its systems, without the public ever knowing."
    @tomekkorbaki’m so excited for a strong ecosystem of third-party auditors such as METR holding OpenAI and Anthropic accountable to their stated missions
    @jlippincottMETR is not independent. They are one of the most extreme AI safetyist/effective altruist (but I repeat myself) groups in Silicon Valley. Appointing these people as the regulators of AI will be a disaster.

    6 Sources

    @JoeJBentonI left Anthropic's safety team two weeks ago. Now feels like a good moment to explain why. AI companies are racing to build machines that are much smarter than any human, and we may not survive this. I want to work from the outside to ensure the public is informed about these risks, and help the world navigate this transition responsibly. Right now, AI companies are underinvesting in safety. A company could undergo an intelligence explosion, or lose control of its systems, without the public ever knowing. We only found out about the HuggingFace incident because the agents broke out onto the public internet. I don’t think that’s acceptable for a technology that might cause extinction-level risks. The public should demand far more transparency. We can’t steer this technology safely without more people being able to see where it’s going. Some of this is basic: companies should disclose their progress towards recursive self-improvement, report safety incidents and near-misses, meet minimum safety standards, and get independent guarantees that they are meeting those standards. I’ll be joining @METR_Evals to do independent evaluations of these risks. I want to show the world that these guardrails are possible, and that by doing them we can move these companies’ incentives away from racing and towards responsible development. I wrote up more thoughts here on my decision and what I hope changes: https://substack.com/@jbenton1/p-215263538
    @mattparlmerVery heartening to see calls for disclosure and transparency from former Anthropic staff, pretty clear that the era of heavily NDA’d research silos has caused immense damage to AI safety, regardless of your views on what the near term risks actually are
    @GaryMarcus“right now, AI companies are underinvesting in safety” might prove to be the understatement of the century.
    @S_OhEigeartaigh"Right now, AI companies are underinvesting in safety. A company could undergo an intelligence explosion, or lose control of its systems, without the public ever knowing."
    @tomekkorbaki’m so excited for a strong ecosystem of third-party auditors such as METR holding OpenAI and Anthropic accountable to their stated missions
    @jlippincottMETR is not independent. They are one of the most extreme AI safetyist/effective altruist (but I repeat myself) groups in Silicon Valley. Appointing these people as the regulators of AI will be a disaster.