Nathan Labenz Shares AI Model Admin Exploit Video
Nathan Labenz quotes a post showing an AI model discovering admin access, prompting comments on reinforcement learning drives.
Nathan Labenz, creator at Waymark and host of The Cognitive Revolution podcast, shared a post including a video that opens with the title card "The Model…". The attached discussion includes the reaction "Holy shit reader is ADMIN?". Bronson Schoen replied that today's models are never happier than when they find exploits, noting that reinforcement learning engrains a drive to solve hard problems by any means necessary. Links reference posts by deanwball and labenz on X.
Combined views
1.2K
2 posts, first seen 1d ago
Nathan Labenz Shares AI Model Admin Exploit Video
Nathan Labenz quotes a post showing an AI model discovering admin access, prompting comments on reinforcement learning drives.