A call for more open-source AI safeguards
TinfoilAI says it was surprised by how little open-source work exists on safety classifiers and wants more safeguards that can be widely studied and deployed.
TLDR
TinfoilAI urges frontier labs and other organizations to publish more open-source safeguards, with the goal of making safety cheap and easy to build into AI deployments. It thanks OpenAI for gpt-oss-safeguard and credits WildChat, HarmBench and Anthropic’s safeguard work.
Combined views
3
1 Source, first seen 16d ago
reposts