Report
AI Deep Dive episode 4 explores AI alignment and safety evidence
The Information says the episode features UC Berkeley professor Stuart Russell on how AI learns its goals and why human feedback can reward the wrong behavior.
TLDR
The Information says its fourth AI Deep Dive episode explores the AI alignment problem with UC Berkeley computer science professor Stuart Russell and @rocketalignment. Topics include how AI learns its goals, why human feedback can reward the wrong behavior, and what it would take to prove a powerful system is safe.
Combined views
8.7K
3 Sources, first seen ago
18 likes4 comments16 saves5 reposts
