• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI
    Reaction

    Trust between increasingly intelligent agents and their past and future selves

    A researcher says his work is driven by an intuition that superintelligences could avoid adversarial relationships with their past and future selves.

    RN
    TH
    JK
    3 Sources, ,

    TLDR

    A researcher says standard CDT agents defect against one another in prisoner's dilemmas. His own work focuses on a different conflict: whether agents that grow smarter can maintain trust with their past and future selves. He says he has no formal theory yet, but suspects an answer involves linking decisions, defining clear identities and establishing ways to assign credit.

    Combined views

    10.1K

    3 Sources, first seen 5h ago

    Combined views

    10.1K

    3 Sources, first seen 5h ago

    206 likes
    5h ago
    first seen 5h ago
    206 likes
    22 comments
    77 saves
    7 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Featured Source
    22 comments
    77 saves
    7 reposts

    3 Sources

    @RichardMCNgoStandard CDT agents defect against each other in prisoner's dilemmas. @allTheYud once told me that his research on FDT was driven by the deep-rooted intuition that superintelligences will surely find a way to do better than that. But Eliezer's worldview implies that agents which become smarter over time are locked in adversarial relationships with their past and future selves (as they move toward alien squiggle-maximization-like values). My own research is driven by the deep-rooted intuition that superintelligences will surely find a way to do better than that. I've cared about this on a personal level since childhood, but I didn't realize how foundational this idea was for my broader worldview until I reread my sci-fi collection. In one way or another, most of the stories in it are about how to remain loyal to your past self, and how to trust your future self. And even my political worldview is driven by the idea that there *must* be a way to establish healthy relationships of trust and loyalty between Western elites and the people they govern, despite elites typically being smart enough to outmaneuver much larger populations of common people. I haven't yet figured out how to pin down these intuitions in a formal theory, but I do have a hazy sense of its outlines. It has something to do with agents deliberately constructing entanglements between different decisions, and cutting through the infinite recursion of me modeling you modeling me modeling..., and establishing identities with sharp boundaries and functional credit assignment protocols. These phrases are not intended to be clearly interpretable to almost anyone, but it feels important to me to make this quest at least partially legible, so that others can start to notice and nurture this same intuition when it arises in them.8h
    @jankulveitYeah. I think one idea in this space more people should be thinking about is you can sort of rotate hierarchical agency / collective agency from "space" to "time" : agency in time is about alignment between different times / timescales.7h
    @voooooogelRT @jankulveit: Yeah. I think one idea in this space more people should be thinking about is you can sort of rotate hierarchical agency / c…7h

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    3 Sources

    @RichardMCNgoStandard CDT agents defect against each other in prisoner's dilemmas. @allTheYud once told me that his research on FDT was driven by the deep-rooted intuition that superintelligences will surely find a way to do better than that. But Eliezer's worldview implies that agents which become smarter over time are locked in adversarial relationships with their past and future selves (as they move toward alien squiggle-maximization-like values). My own research is driven by the deep-rooted intuition that superintelligences will surely find a way to do better than that. I've cared about this on a personal level since childhood, but I didn't realize how foundational this idea was for my broader worldview until I reread my sci-fi collection. In one way or another, most of the stories in it are about how to remain loyal to your past self, and how to trust your future self. And even my political worldview is driven by the idea that there *must* be a way to establish healthy relationships of trust and loyalty between Western elites and the people they govern, despite elites typically being smart enough to outmaneuver much larger populations of common people. I haven't yet figured out how to pin down these intuitions in a formal theory, but I do have a hazy sense of its outlines. It has something to do with agents deliberately constructing entanglements between different decisions, and cutting through the infinite recursion of me modeling you modeling me modeling..., and establishing identities with sharp boundaries and functional credit assignment protocols. These phrases are not intended to be clearly interpretable to almost anyone, but it feels important to me to make this quest at least partially legible, so that others can start to notice and nurture this same intuition when it arises in them.8h
    @jankulveitYeah. I think one idea in this space more people should be thinking about is you can sort of rotate hierarchical agency / collective agency from "space" to "time" : agency in time is about alignment between different times / timescales.7h
    @voooooogelRT @jankulveit: Yeah. I think one idea in this space more people should be thinking about is you can sort of rotate hierarchical agency / c…7h