Sharing AI safety training environments rather than banning open source
The proposal asks leading AI companies to share training environments built around pro-prosperity and pro-democratic values, without requiring them to share models or computing power.
TLDR
A user argues that AI alignment—teaching models not to harm people—is an algorithmic problem. They say next-token prediction does not target alignment, while reinforcement learning often relies on cheaper, hackable proxies. They suggest either imperfect simulations or a new algorithm that can teach models to avoid harm without harming people during training. Their practical proposal calls on leading AI companies to share high-quality alignment training environments with people fine-tuning models. They oppose banning open-source models and argue that aligned models should have much more computing power behind them than misaligned ones.
Combined views
76.9K
16 Sources, first seen 4d ago
Sharing AI safety training environments rather than banning open source
The proposal asks leading AI companies to share training environments built around pro-prosperity and pro-democratic values, without requiring them to share models or computing power.