• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    The push for more open-source AI safeguards

    TinfoilAI urges frontier labs and other organizations to publish more safeguards that can be widely studied and deployed, with the goal of making safety cheap and easy to build into AI systems.

    TI
    1 Source, 16d ago, first seen 16d ago

    TLDR

    TinfoilAI says it was surprised by how little open-source work exists on safety classifiers and calls for frontier labs and other organizations to publish more open-source safeguards. It hopes to contribute to that ecosystem. The organization also thanks OpenAI for gpt-oss-safeguard, the creators of WildChat and HarmBench, and Anthropic for publishing safeguard work.

    Combined views

    480

    1 Source, first seen 16d ago

    Combined views

    480

    1 Source, first seen 16d ago

    11 likes
    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    11 likes
    1 comments
    1 saves
    2 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    1 comments
    1 saves
    2 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    1 Source

    @TinfoilAIThanks to @OpenAI for gpt-oss-safeguard, @yuntiandeng for WildChat, @MantasMazeika96 + @CAIS for HarmBench and @AnthropicAI for publish safeguard work. We were surprised by how little open-source work exists on safety classifiers. We urge frontier labs and other organizations to publish more open-source safeguards that can be widely studied and deployed, making it cheap and easy to build safety into any AI deployment. We're excited to see this ecosystem grow and hope to contribute.

    1 Source

    @TinfoilAIThanks to @OpenAI for gpt-oss-safeguard, @yuntiandeng for WildChat, @MantasMazeika96 + @CAIS for HarmBench and @AnthropicAI for publish safeguard work. We were surprised by how little open-source work exists on safety classifiers. We urge frontier labs and other organizations to publish more open-source safeguards that can be widely studied and deployed, making it cheap and easy to build safety into any AI deployment. We're excited to see this ecosystem grow and hope to contribute.