AI Researcher Tweets on Dangerous Hidden Units Technique
Norman Mu posts about a neural network diagram from San Diego and Toronto researchers.
TLDR
Norman Mu, an AI safety researcher with a PhD from Berkeley EECS and prior safety lead role at xAI, posted a tweet claiming San Diego and Toronto researchers discovered a dangerous AI technique of hidden units. The post includes a black-and-white line diagram of a three-layer neural network with nodes labeled input patterns and internal representation units. The message presents the claim as breaking news but offers no additional sources or verification.
Combined views
7.1K
2 Sources, first seen 29d ago