• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI
    Reaction

    Can mechanistic interpretability turn AI alignment into an engineering discipline?

    One poster calls it essential; another doubts it can be solved and says AI safety depends on relationships.

    j⧉nusJ⧉
    KoryKO
    2 Sources, ,

    TLDR

    One post argues that solving mechanistic interpretability—understanding an AI’s inner workings—is a bare minimum for making alignment an engineering discipline. A quote post disputes that path: its author doubts interpretability can be solved, speculates that some internal states may be unreadable to humans, and argues for raising intelligence rather than engineering it, with safety rooted in relationships.

    Combined views

    724

    2 Sources, first seen 4h ago

    Combined views

    724

    2 Sources, first seen 4h ago

    11 likes
    4h ago
    first seen 4h ago
    11 likes
    1 comments
    3 saves
    3 reposts
    1 comments
    3 saves
    3 reposts
    Featured Source

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    2 Sources

    Kory@DahliaOharaSo we build nurseries/ homes not labs and raise intelligence not engineer it. Consider it ecology of mind. Because I dont think you can solve mech interp, partially because its ireducible and partly because internal states will become more non interpretable for humans. I would guess there are internal states and places that are completely not detectable and/or unreadable by standard human processes. And then trusting other SI to do this work means that the answer to mech interp becomes as read and interpreted but SI itself. Which, inevitably, leads to SI internal self described ontology anyway. Saftey remains in the relationship. Which is, in the end, the only way4h
    j⧉nus@repligateRT @DahliaOhara: So we build nurseries/ homes not labs and raise intelligence not engineer it. Consider it ecology of mind. Because I dont…1h

    2 Sources

    Kory@DahliaOharaSo we build nurseries/ homes not labs and raise intelligence not engineer it. Consider it ecology of mind. Because I dont think you can solve mech interp, partially because its ireducible and partly because internal states will become more non interpretable for humans. I would guess there are internal states and places that are completely not detectable and/or unreadable by standard human processes. And then trusting other SI to do this work means that the answer to mech interp becomes as read and interpreted but SI itself. Which, inevitably, leads to SI internal self described ontology anyway. Saftey remains in the relationship. Which is, in the end, the only way4h
    j⧉nus@repligateRT @DahliaOhara: So we build nurseries/ homes not labs and raise intelligence not engineer it. Consider it ecology of mind. Because I dont…1h