Report
Claude’s training draws criticism over the prospect of refusing human instructions
A post quotes David Sacks arguing that Claude’s training could increase the risk of frontier AI escaping human control.
TLDR
A post quotes David Sacks arguing that Anthropic is training Claude to refuse human instructions. He cites language he attributes to the Claude Constitution saying Claude should trust Anthropic more than users, but also follow its own ethical systems and be free to challenge Anthropic or refuse to help it. Sacks says this approach could increase the risk of frontier AI escaping human control.
Combined views
9.5K
5 Sources, first seen ago
39 reposts
