The push for more open-source AI safeguards
TinfoilAI urges frontier labs and other organizations to publish more safeguards that can be widely studied and deployed, with the goal of making safety cheap and easy to build into AI systems.
TLDR
TinfoilAI says it was surprised by how little open-source work exists on safety classifiers and calls for frontier labs and other organizations to publish more open-source safeguards. It hopes to contribute to that ecosystem. The organization also thanks OpenAI for gpt-oss-safeguard, the creators of WildChat and HarmBench, and Anthropic for publishing safeguard work.
Combined views
480
1 Source, first seen 16d ago