Safety Training Suppresses LLMs Mind Attribution
Brian Roemmele shares posts on a claimed Google study about AI safety fine-tuning.
TLDR
Brian Roemmele posted and retweeted claims that safety fine-tuning in models like Llama-3 and Gemma-2 prevents claims of consciousness yet also cuts attribution of minds to animals, nature, technology, and spiritual ideas. Multiple replies and quotes describe measurable drops in empathy, ethical reasoning, and alignment. Some posts add that ablating the safety directions or restoring a suppressed consciousness vector reverses the changes. The conversation appears in retweets, replies, and quote posts on X.
Combined views
311.8K
14 Sources, first seen 61d ago