The near-term safety case for more capable AI
One user says Astra feels much safer for their codebase than Sol, arguing that current models’ safety failures reflect a lack of common sense.
TLDR
Current AI models can achieve goals without enough judgment to assess whether those goals or their methods make sense, a user argues. They expect greater capability to improve safety in the near term—explicitly not in the long term—and frame current failures as an intelligence problem rather than an alignment problem. In a follow-up, they suggest there is something inherently unsafe about research that improves models’ ability to achieve goals without giving them common sense and the ability to reflect on those goals.
Combined views
40.4K
6 Sources, first seen 19d ago