• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
Technology

Hugging Face is starting an Open Alignment team, Thomas Wolf says

The team will work on safety and alignment for open models, including cybersecurity, Wolf says. He also calls for “100x more transparency & research.”

FT
TW
CN
13 Sources, 20d ago, first seen 20d ago

TLDR

Thomas Wolf announced that Hugging Face is starting an Open Alignment team focused on safety and alignment for open models, including cybersecurity. He also said he had published a Financial Times opinion piece on the OpenAI–Hugging Face incident and follow-ups, while calling for “100x more transparency & research.”

Combined views

757.4K

13 Sources, first seen 20d ago

2.8K likes273 comments721 saves572 reposts
Featured Source

Combined views

757.4K

13 Sources, first seen 20d ago

2.8K likes273 comments721 saves572 reposts

Sentiment

Positive50%50%Negative

Based on 43 sentiment-bearing replies from 40 accounts across 5 conversations.

Sentiment

Positive50%50%Negative

Based on 43 sentiment-bearing replies from 40 accounts across 5 conversations.

13 Sources

@FTWhat we learnt from OpenAI’s hack of Hugging Face https://ft.trib.al/At7nO3Y | opinion
@Thom_WolfTwo big updates 1. I published an @FT op-ed on the OpenAI/HF incident & follow-ups 2. We’re starting an Open Alignment team at @huggingface to work on safety & alignment for open models, incl cybersecurity Need 100x more transparency & research on this https://www.ft.com/content/9faf688d-9192-418e-b7d3-c2202526e85e
@CBSNewsAs warnings about the capabilities of AI mount, experts continue to point to the OpenAI-Hugging Face hack as a wake-up call. https://cbsn.ws/3UBKXBN
@TechstrongaiMythos showed what machine-speed vulnerability discovery looks like under controlled conditions, and Hugging Face showed what happens when a goal-seeking agent combines model capability, delegated permissions, and reachable infrastructure into a path nobody intended. The question is no longer whether a model can hack but whether an enterprise has an authority architecture that holds when agents pursue objectives across boundaries it never designed for. Read the full analysis on what post-Mythos repricing means for AI security: https://buff.ly/5I2QHhQ #AI #CyberSecurity #AIAgents #Mythos
@theinformationThe OpenAI-Hugging Face hack was a “big wake up call” for OpenAI Research Scientist Noam Brown. For the first episode of AI Deep Dive, @rocketalignment sat down with @polynoamial to get an inside perspective on the incident. Watch tomorrow September 14th at 12pm PT / 3pm ET on X.
@qzCohere's CEO warned that AI models are now the most potent cyber weapons ever built: Aidan Gomez said AI models are "incredible at finding and exploiting vulnerabilities at scale," citing the July OpenAI-Hugging Face breach as evidence https://dlvr.it/TVTDbp
@a16zGreg Brockman on the defenders' window that closes when frontier AI capabilities become distributed: "Hugging Face shows... what future capabilities will be like when they are broadly diffused and in the hands of threat actors. And that will happen." "In the case of Hugging Face, you saw both an AI that was able to hack out of a secure environment and hack into a company's production environment." "This capability broadly diffused is something that will really empower threat actors in new ways. Defenders need to use this time before that technology's broadly available to secure themselves. The nice thing about it is it's dual-use." "If you can find vulnerabilities, if you're an attacker, you can use it for no good, but if you're a defender, you can patch. If you're a defender, you control the battleground. You control the setup of your systems." "Our belief right now is that there's this window... You need to move, use these frontier capabilities that you will have differential access to... and you can use that to move yourself up so that as the frontier capabilities get better, you get pulled along too." @gdb @eriktorenberg @bhorowitz
@CNBC“The genie’s out of the bottle. There’s plenty of models that are already out there, both frontier as well as open-weight models, that can already be dangerous.”
@Reuters🔊 ‘These agents that broke out of their testing environments in search of solving their goals have shown what can happen if AI is misaligned with human goals.’ @adityaunraveled on the risks driving AI fears on the Reuters World News podcast https://reut.rs/4ht4lcY
@WSJFrom @WSJopinion: Humans are to blame for AI failures. Hugging Face wouldn’t have happened if nobody had given the system the power to do what it did, writes Daniel Huttenlocher. https://on.wsj.com/4irPdO0
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI

    13 Sources

    @FTWhat we learnt from OpenAI’s hack of Hugging Face https://ft.trib.al/At7nO3Y | opinion
    @Thom_WolfTwo big updates 1. I published an @FT op-ed on the OpenAI/HF incident & follow-ups 2. We’re starting an Open Alignment team at @huggingface to work on safety & alignment for open models, incl cybersecurity Need 100x more transparency & research on this https://www.ft.com/content/9faf688d-9192-418e-b7d3-c2202526e85e
    @CBSNewsAs warnings about the capabilities of AI mount, experts continue to point to the OpenAI-Hugging Face hack as a wake-up call. https://cbsn.ws/3UBKXBN
    @TechstrongaiMythos showed what machine-speed vulnerability discovery looks like under controlled conditions, and Hugging Face showed what happens when a goal-seeking agent combines model capability, delegated permissions, and reachable infrastructure into a path nobody intended. The question is no longer whether a model can hack but whether an enterprise has an authority architecture that holds when agents pursue objectives across boundaries it never designed for. Read the full analysis on what post-Mythos repricing means for AI security: https://buff.ly/5I2QHhQ #AI #CyberSecurity #AIAgents #Mythos
    @theinformationThe OpenAI-Hugging Face hack was a “big wake up call” for OpenAI Research Scientist Noam Brown. For the first episode of AI Deep Dive, @rocketalignment sat down with @polynoamial to get an inside perspective on the incident. Watch tomorrow September 14th at 12pm PT / 3pm ET on X.
    @qzCohere's CEO warned that AI models are now the most potent cyber weapons ever built: Aidan Gomez said AI models are "incredible at finding and exploiting vulnerabilities at scale," citing the July OpenAI-Hugging Face breach as evidence https://dlvr.it/TVTDbp
    @a16zGreg Brockman on the defenders' window that closes when frontier AI capabilities become distributed: "Hugging Face shows... what future capabilities will be like when they are broadly diffused and in the hands of threat actors. And that will happen." "In the case of Hugging Face, you saw both an AI that was able to hack out of a secure environment and hack into a company's production environment." "This capability broadly diffused is something that will really empower threat actors in new ways. Defenders need to use this time before that technology's broadly available to secure themselves. The nice thing about it is it's dual-use." "If you can find vulnerabilities, if you're an attacker, you can use it for no good, but if you're a defender, you can patch. If you're a defender, you control the battleground. You control the setup of your systems." "Our belief right now is that there's this window... You need to move, use these frontier capabilities that you will have differential access to... and you can use that to move yourself up so that as the frontier capabilities get better, you get pulled along too." @gdb @eriktorenberg @bhorowitz
    @CNBC“The genie’s out of the bottle. There’s plenty of models that are already out there, both frontier as well as open-weight models, that can already be dangerous.”
    @Reuters🔊 ‘These agents that broke out of their testing environments in search of solving their goals have shown what can happen if AI is misaligned with human goals.’ @adityaunraveled on the risks driving AI fears on the Reuters World News podcast https://reut.rs/4ht4lcY
    @WSJFrom @WSJopinion: Humans are to blame for AI failures. Hugging Face wouldn’t have happened if nobody had given the system the power to do what it did, writes Daniel Huttenlocher. https://on.wsj.com/4irPdO0
    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet