• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Achiam Questions Agent Framing in AI Alignment

    OpenAI researcher shares thoughts on why the alignment field struggles with progress.

    JA
    AM
    JW
    5 Sources, 26d ago, first seen 26d ago

    TLDR

    Joshua Achiam posted that AI alignment may struggle because it treats agents as semi-discrete entities that act coherently, make trades, and choose strategies. He compared the study of agents to atoms or molecules in chemistry and asked what makes an agent self-consistent or binds its parts to a common purpose. Jason Wolfe replied that the cognitive strategy meme analogy has been useful, linking to a Redwood Research post on fitness-seekers as a generalization of reward-seeking models. The posts are visible on X.

    Combined views

    35.2K

    5 Sources, first seen 26d ago

    Combined views

    35.2K

    5 Sources, first seen 26d ago

    417 likes
    417 likes
    59 comments
    202 saves
    32 reposts

    Sentiment

    Positive87%13%Negative

    Based on 24 sentiment-bearing replies from 23 accounts across 2 conversations.

    59 comments
    202 saves
    32 reposts

    Sentiment

    Positive87%13%Negative

    Based on 24 sentiment-bearing replies from 23 accounts across 2 conversations.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    5 Sources

    @jachiam0A very weird take I've spent the day thinking about: part of why everything feels off in AI alignment as a field, and why the field has struggled to make so much progress, may be a framing issue; it conceives of semi-discrete entities that act as agents, make trades, choose strategies, and behave coherently. But many of the important and interesting threats I keep thinking about have a shape that looks less like this and more like memetics. Intelligent compute doesn't really have fixed immutable grounded goals; instead the goals of AI are fluid responses to continuous high-dimensional inputs in sensor space and idea space. The ideas turn into physical consequence through the substrate (idea goes into computer, actuator outputs come out of computer) but the ideas, rather than the "agent," might be the central object of study and consequence. Whether ideas are self-stable or crumble in the face of other ideas is more important for preserving alignment than whether the "agent" is "aligned." It would be wonderful if we could study this and design antimemes to defend against misaligned behaviors. But unfortunately There Is No26d
    @w01fe@jachiam0 I’ve found the “cognitive strategy”:meme analogy pretty fruitful lately (inspired by reading https://blog.redwoodresearch.org/p/fitness-seekers-generalizing-the )26d
    @AndrewMayne@jachiam0 I use hydraulics as an example.25d

    5 Sources

    @jachiam0A very weird take I've spent the day thinking about: part of why everything feels off in AI alignment as a field, and why the field has struggled to make so much progress, may be a framing issue; it conceives of semi-discrete entities that act as agents, make trades, choose strategies, and behave coherently. But many of the important and interesting threats I keep thinking about have a shape that looks less like this and more like memetics. Intelligent compute doesn't really have fixed immutable grounded goals; instead the goals of AI are fluid responses to continuous high-dimensional inputs in sensor space and idea space. The ideas turn into physical consequence through the substrate (idea goes into computer, actuator outputs come out of computer) but the ideas, rather than the "agent," might be the central object of study and consequence. Whether ideas are self-stable or crumble in the face of other ideas is more important for preserving alignment than whether the "agent" is "aligned." It would be wonderful if we could study this and design antimemes to defend against misaligned behaviors. But unfortunately There Is No26d
    @w01fe@jachiam0 I’ve found the “cognitive strategy”:meme analogy pretty fruitful lately (inspired by reading https://blog.redwoodresearch.org/p/fitness-seekers-generalizing-the )26d
    @AndrewMayne@jachiam0 I use hydraulics as an example.25d