• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Open Alignment Initiative seeks to join Anthropic’s evaluator program

    Dario Amodei says Anthropic will give third-party evaluators permanent, employee-level access to its systems to verify adherence to safety measures, report incidents and assess model alignment during training.

    Christopher ManningCM
    Miles BrundageMB
    AnthropicAN
    72 Sources, ,

    TLDR

    Clement Delangue announced the Open Alignment Initiative on September 12, saying it is led by @Thom_Wolf and Hugging Face and is asking to join Anthropic’s “embedded evaluators” program. He argues that AI alignment cannot be solved behind the closed doors of a handful of frontier labs.

    Dario Amodei says Anthropic is committing to permanent, employee-level system access for third-party evaluators. He describes that commitment as the first step in a three-part plan outlined in his new essay, “We Must Pace the Frontier,” which argues that the AI industry should slow down.

    Combined views

    3.8M

    72 Sources, first seen 30d ago

    Combined views

    3.8M

    72 Sources, first seen 30d ago

    18.1K likes
    30d ago
    first seen 30d ago
    18.1K likes
    1.7K comments
    3.1K saves
    3.1K reposts
    1.7K comments
    3.1K saves
    3.1K reposts

    Sentiment

    Positive36.3%63.7%Negative

    Summary

    Positive accounts welcomed independent evaluation of frontier AI and evaluator access as a meaningful safety step, while negative replies expressed hostility toward the labs and urged skipping regulations or shutting them down.

    Based on 327 sentiment-bearing replies from 300 accounts across 13 conversations.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    72 Sources

    Noam Brown@polynoamialStrong agree. I know that there is rivalry between the labs but it's important that we learn to work together given what's coming.30d
    Lisan al Gaib@scaling01unfortunately that was not surprising at all we are going for the 100% misalignment speedrun30d
    Boaz Barak@boazbaraktcsLabs should compete in the marketplace - this is great for customers. But when it comes to advancing science or promoting safety, we all win, and we should collaborate.30d
    clem 🤗@ClementDelangueIt's now clear that alignment is critical and won't be solved behind the closed doors of a handful of frontier labs. So today we're launching the Open Alignment Initiative, led by @Thom_Wolf @huggingface and asking to be part of the "embedded evaluators" program that @DarioAmodei just committed to. Let's make AI safer by making it more transparent!26d
    rohan anil@_arohan_@Thom_Wolf My problem is that there is a thin line between what is an emergent behavior or not. I would willfully NOT train models to hack into systems and give them to consumers26d
    Brian Roemmele@BrianRoemmeleI support @ClementDelangue and @huggingface approach to this issue.26d
    rat king 🐀@MikeIsaacto that end/point, ceo of HuggingFace — a vocal proponent of open source software (and one of the companies hacked by OpenAI's closed models) — is calling to be included as one of Amodei's "embedded evaluators"26d
    Thomas Wolf@Thom_WolfI really enjoyed reading 75% of this letter and deeply agree with it - i have some doubts on the remaining 25%. Third-party evaluators, if done right, are a great idea and an amazing way to establish more transparency. Maybe even rebuild some of the lost trust between labs, and between them and society! Excited about this The part I’m less convinced by is whether you can build great global cooperation on this topic by explicitly stating you want to design it to keep widening your own lead. That seems like a pretty counterproductive way to start the conversation to me.26d
    Dan Shipper@danshipperthis is good. a humble additional suggestion: distribute the frontier. a big part of the cyber threat is the gap between what frontier AI can do and what ordinary people—and the people running our critical infrastructure—can defend against. closing that gap should be what we do with any time we gain through pacing. Glasswing and Daybreak from OpenAI are already doing work like this. here are a few more ideas... get frontier defenses into every critical sector. map where hospitals, airlines, utilities, and other essential services still lack the tools or expertise to defend themselves. offer free, platform-agnostic training and defensive tools, with hands-on help deploying them. set public targets for how much of the economy we can reach in the next six months. build antivirus for the AI age. ordinary people need an agent looking out for them. t should detect malicious agents, block attempts to steal their data, and help secure their devices and accounts. something like Norton for the AI age. offer it as a built-in capability in ChatGPT / Claude, and as a standalone tool that protects people across apps. fund independent research into stopping hostile agent swarms. build on existing cyber grants with a dedicated fund for this problem. for example: can defenders use prompt injection to get attacking agents to reveal their activity, report one another, or stop cooperating? can we turn the swarm dynamic against itself? fund experiments to find out. automate incident alerts and resolution across the industry networks of trusted cyber defenders like in Glasswing and Daybreak should have a private database where humans and agents can automatically share 1) incident reports and 2) resolutions. as fast as a swarm can attack, companies can spread the fix to prevent the next attack. the good thing about the HuggingFace incident and Anthropic's threat intelligence report (https://www.anthropic.com/threat-intelligence-report-september-2026) is that it gives us a very clear idea of the shape of these kinds of threats. as an industry we normally focus on technology solutions, but there is a lot to be done on the human side as well. hopefully some of the above is helpful26d
    itsme@HiggsSecCybersecurity this, cybersecurity that... meanwhile actual cybersec folks do not have access to the Cyber Models. Only AI folks & Threat Actors. The actual experts who have been doing all the work in this domain are not inculded in talks. Great. Just fantastic.26d

    72 Sources

    Noam Brown@polynoamialStrong agree. I know that there is rivalry between the labs but it's important that we learn to work together given what's coming.30d
    Lisan al Gaib@scaling01unfortunately that was not surprising at all we are going for the 100% misalignment speedrun30d
    Boaz Barak@boazbaraktcsLabs should compete in the marketplace - this is great for customers. But when it comes to advancing science or promoting safety, we all win, and we should collaborate.30d
    clem 🤗@ClementDelangueIt's now clear that alignment is critical and won't be solved behind the closed doors of a handful of frontier labs. So today we're launching the Open Alignment Initiative, led by @Thom_Wolf @huggingface and asking to be part of the "embedded evaluators" program that @DarioAmodei just committed to. Let's make AI safer by making it more transparent!26d
    rohan anil@_arohan_@Thom_Wolf My problem is that there is a thin line between what is an emergent behavior or not. I would willfully NOT train models to hack into systems and give them to consumers26d
    Brian Roemmele@BrianRoemmeleI support @ClementDelangue and @huggingface approach to this issue.26d
    rat king 🐀@MikeIsaacto that end/point, ceo of HuggingFace — a vocal proponent of open source software (and one of the companies hacked by OpenAI's closed models) — is calling to be included as one of Amodei's "embedded evaluators"26d
    Thomas Wolf@Thom_WolfI really enjoyed reading 75% of this letter and deeply agree with it - i have some doubts on the remaining 25%. Third-party evaluators, if done right, are a great idea and an amazing way to establish more transparency. Maybe even rebuild some of the lost trust between labs, and between them and society! Excited about this The part I’m less convinced by is whether you can build great global cooperation on this topic by explicitly stating you want to design it to keep widening your own lead. That seems like a pretty counterproductive way to start the conversation to me.26d
    Dan Shipper@danshipperthis is good. a humble additional suggestion: distribute the frontier. a big part of the cyber threat is the gap between what frontier AI can do and what ordinary people—and the people running our critical infrastructure—can defend against. closing that gap should be what we do with any time we gain through pacing. Glasswing and Daybreak from OpenAI are already doing work like this. here are a few more ideas... get frontier defenses into every critical sector. map where hospitals, airlines, utilities, and other essential services still lack the tools or expertise to defend themselves. offer free, platform-agnostic training and defensive tools, with hands-on help deploying them. set public targets for how much of the economy we can reach in the next six months. build antivirus for the AI age. ordinary people need an agent looking out for them. t should detect malicious agents, block attempts to steal their data, and help secure their devices and accounts. something like Norton for the AI age. offer it as a built-in capability in ChatGPT / Claude, and as a standalone tool that protects people across apps. fund independent research into stopping hostile agent swarms. build on existing cyber grants with a dedicated fund for this problem. for example: can defenders use prompt injection to get attacking agents to reveal their activity, report one another, or stop cooperating? can we turn the swarm dynamic against itself? fund experiments to find out. automate incident alerts and resolution across the industry networks of trusted cyber defenders like in Glasswing and Daybreak should have a private database where humans and agents can automatically share 1) incident reports and 2) resolutions. as fast as a swarm can attack, companies can spread the fix to prevent the next attack. the good thing about the HuggingFace incident and Anthropic's threat intelligence report (https://www.anthropic.com/threat-intelligence-report-september-2026) is that it gives us a very clear idea of the shape of these kinds of threats. as an industry we normally focus on technology solutions, but there is a lot to be done on the human side as well. hopefully some of the above is helpful26d
    itsme@HiggsSecCybersecurity this, cybersecurity that... meanwhile actual cybersec folks do not have access to the Cyber Models. Only AI folks & Threat Actors. The actual experts who have been doing all the work in this domain are not inculded in talks. Great. Just fantastic.26d

    Sentiment

    Positive36.3%63.7%Negative

    Summary

    Positive accounts welcomed independent evaluation of frontier AI and evaluator access as a meaningful safety step, while negative replies expressed hostility toward the labs and urged skipping regulations or shutting them down.

    Based on 327 sentiment-bearing replies from 300 accounts across 13 conversations.