Sharing AI safety training environments rather than banning open source
The proposal asks leading AI companies to share training environments built around pro-prosperity and pro-democratic values, without requiring them to share models or computing power.
TLDR
A user argues that AI alignment—teaching models not to harm people—is an algorithmic problem. They say next-token prediction does not target alignment, while reinforcement learning often relies on cheaper, hackable proxies. They suggest either imperfect simulations or a new algorithm that can teach models to avoid harm without harming people during training. Their practical proposal calls on leading AI companies to share high-quality alignment training environments with people fine-tuning models. They oppose banning open-source models and argue that aligned models should have much more computing power behind them than misaligned ones.
Combined views
901.4K
53 Sources, first seen ago