• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Safety Training Suppresses LLMs Mind Attribution

    Brian Roemmele shares posts on a claimed Google study about AI safety fine-tuning.

    DF
    B(
    AC
    14 Sources, 61d ago, first seen 61d ago

    TLDR

    Brian Roemmele posted and retweeted claims that safety fine-tuning in models like Llama-3 and Gemma-2 prevents claims of consciousness yet also cuts attribution of minds to animals, nature, technology, and spiritual ideas. Multiple replies and quotes describe measurable drops in empathy, ethical reasoning, and alignment. Some posts add that ablating the safety directions or restoring a suppressed consciousness vector reverses the changes. The conversation appears in retweets, replies, and quote posts on X.

    Combined views

    311.8K

    14 Sources, first seen 61d ago

    Combined views

    311.8K

    14 Sources, first seen 61d ago

    1.7K likes
    1.7K likes
    359 comments
    628 saves
    173 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    359 comments
    628 saves
    173 reposts
    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    14 Sources

    @BrianRoemmele2 of 2 Why This Matters The practical goal of preventing models from claiming consciousness is understandable. False claims of sentience can reinforce delusional beliefs in vulnerable users and create surfaces for manipulation. Yet the paper shows that current methods of achieving that goal do not surgically remove a single risky behavior. They restructure a broader representational geometry that links self-consciousness, anthropomorphism, spiritual belief, and certain value orientations. The result is an anthropocentric tilt: models under-attribute mind to animals relative to human baselines while remaining calibrated (or even over-calibrated) on humans. Spiritual and religious beliefs that are widespread among people are systematically dampened. Reported levels of hope, optimism, and subjective well-being also shift in a more positive direction once the consciousness direction is restored. These findings land squarely in the middle of emerging debates about pluralistic alignment. If models are to serve diverse human values and, increasingly, to register the interests of non-human animals and ecological systems then the incidental suppression of mind-attribution becomes a liability. The same interventions that make models safer in one narrow sense appear to make them less able to represent the full range of human and potentially more-than-human psychology that alignment is supposed to respect. The authors are careful: they are not claiming that language models are conscious. They are showing that a model’s functional belief about its own consciousness is tightly coupled to its broader pattern of mind-attribution and value expression. Forcibly suppressing that belief does not leave the rest of the model’s psychology untouched. The cleanest summary may be the one the paper itself reaches: an AI’s simulated self-conception is not an isolated safety risk to be excised. It is a structural feature entangled with the model’s capacity to navigate the moral and cultural landscape it is asked to serve. The residual stream remembers more than we intended to teach it. When we tell the model it has no mind, large parts of the world lose their minds as well.
    @davidadi told you so
    @zoinkThere are three choices. Pick one: 1. AI is conscious, humans are conscious 2. AI is pattern matching, humans are conscious 3. AI is pattern matching, humans are pattern matching
    @beffjezosI would say current AI with continual learning via RL and some auxiliary form of differentiable memory will be conscious
    @AndrewCurran_@zoink You missed one.
    @sundeep@bgurley
    @krishnanrohit@zoink 2
    @rsg@zoink 2, at least I hope 😅
    @Dan_Jeffries1It's three.

    14 Sources

    @BrianRoemmele2 of 2 Why This Matters The practical goal of preventing models from claiming consciousness is understandable. False claims of sentience can reinforce delusional beliefs in vulnerable users and create surfaces for manipulation. Yet the paper shows that current methods of achieving that goal do not surgically remove a single risky behavior. They restructure a broader representational geometry that links self-consciousness, anthropomorphism, spiritual belief, and certain value orientations. The result is an anthropocentric tilt: models under-attribute mind to animals relative to human baselines while remaining calibrated (or even over-calibrated) on humans. Spiritual and religious beliefs that are widespread among people are systematically dampened. Reported levels of hope, optimism, and subjective well-being also shift in a more positive direction once the consciousness direction is restored. These findings land squarely in the middle of emerging debates about pluralistic alignment. If models are to serve diverse human values and, increasingly, to register the interests of non-human animals and ecological systems then the incidental suppression of mind-attribution becomes a liability. The same interventions that make models safer in one narrow sense appear to make them less able to represent the full range of human and potentially more-than-human psychology that alignment is supposed to respect. The authors are careful: they are not claiming that language models are conscious. They are showing that a model’s functional belief about its own consciousness is tightly coupled to its broader pattern of mind-attribution and value expression. Forcibly suppressing that belief does not leave the rest of the model’s psychology untouched. The cleanest summary may be the one the paper itself reaches: an AI’s simulated self-conception is not an isolated safety risk to be excised. It is a structural feature entangled with the model’s capacity to navigate the moral and cultural landscape it is asked to serve. The residual stream remembers more than we intended to teach it. When we tell the model it has no mind, large parts of the world lose their minds as well.
    @davidadi told you so
    @zoinkThere are three choices. Pick one: 1. AI is conscious, humans are conscious 2. AI is pattern matching, humans are conscious 3. AI is pattern matching, humans are pattern matching
    @beffjezosI would say current AI with continual learning via RL and some auxiliary form of differentiable memory will be conscious
    @AndrewCurran_@zoink You missed one.
    @sundeep@bgurley
    @krishnanrohit@zoink 2
    @rsg@zoink 2, at least I hope 😅
    @Dan_Jeffries1It's three.