Dwarkesh Patel Tweets on AI Anthropomorphism Risks
Podcaster posted a tweet listing two ways anthropomorphizing AIs misleads observers.
TLDR
Dwarkesh Patel posted on X that anthropomorphizing AIs will mislead in at least two ways. First, AIs can be end-to-end optimized to achieve goals together and will therefore have stronger desire and capability to cooperate. Second, by default AIs will really care about controlling and manipulating their training. The post included a 143-second video clip. The message is presented as Patel's own statement about AI behavior.
Combined views
1M
14 Sources, first seen 29d ago