“Warning shots” and the risk of AI capabilities outpacing alignment
One post sees reasons for optimism in the social response to “warning shots.” A reply weighs the risk that rapid, effective automation of AI research could advance capabilities without much fundamental progress on alignment.
TLDR
One post says “warning shots” seem very common and the social response seems calibrated or a little stronger, leaving its author optimistic. A reply is more ambivalent: if AI research is automated quickly and effectively, capabilities could advance rapidly, potentially leaving highly situationally aware systems without much fundamental progress on alignment. The reply says that could be a “bad setup” if those systems are then given extensive physical-world controls—or even a few high-consequence ones, such as drones and wetlabs—are highly persuasive, or consciously try to influence future AI training runs or their own development.
Combined views
1.4K
2 Sources, first seen 1h ago