• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    DeepMind Economist Shares Mechanism Design Framework for AI Safety

    Alex Imas announces a new hire and highlights mechanism design work for multi-agent AI safety.

    RL
    SK
    D🎇
    10 Sources, 28d ago, first seen 28d ago

    TLDR

    Alex Imas, Director of AGI Economics at Google DeepMind, posted that the newest member of the Economics team sees mechanism design as a key tool for alignment and safety with multi-agent systems. He pointed to a framework presented by Andrew, Dirk, and Stephen for safety practices. William Isaac, Principal Scientist on Google DeepMind's Ethics and Society Team, retweeted the post. The statement comes directly from Imas about team priorities.

    Combined views

    169.2K

    10 Sources, first seen 28d ago

    Combined views

    169.2K

    10 Sources, first seen 28d ago

    1K likes
    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    1K likes
    30 comments
    784 saves
    326 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    30 comments
    784 saves
    326 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    10 Sources

    @andrewjkohWe develop a mechanism design framework for AI alignment and control: https://arxiv.org/abs/2609.01595 It’s largely conceptual but we offer stylized applications to failure modes (sandbagging, alignment faking), safety practice (scalable oversight, peer prediction), and a way to think about the value of alignment, interpretability, capability, and control.
    @alexolegimasFrom the newest member of our Economics team at @GoogleDeepMind: Mechanism design will be a key tool for alignment and safety with multi-agent systems. Andrew, Dirk, and Stephen present a rigorous framework for how to think about safety practices from a mechanism design perspective.
    @wsisaacRT @alexolegimas: From the newest member of our Economics team at @GoogleDeepMind: Mechanism design will be a key tool for alignment and sa…
    @malleshpaiReally nice paper!! There's a lovely line of work forming, this paper included, on applying mechanism design to AI alignment. AIs are autonomous black boxes that may well end up with preferences of their own. Alignment research mostly asks: how do we shape those preferences? But as models get more capable, it's getting harder and harder to directly shape those preferences and/or have confidence that we are shaping them as desired. Mechanism design asks the complementary question economists have asked about humans for decades: taking preferences as unknown and possibly bad, how do we design the rules (evals, permissions, rewards) so we get good outcomes anyway? We should be doing much more of the second.
    @sebkrierRT @andrewjkoh: We develop a mechanism design framework for AI alignment and control: https://arxiv.org/abs/2609.01595 It’s largely conceptual but…
    @ryan_t_lowedang, didn't realize that @andrewjkoh had joined DeepMind, great get
    @davidadThis is really cool and important work. Conjecture: this mechanism and the self-DPO I’ve been on about are instances of a common principle—to get a mind to fall into a robust basin, an RL scorer must use the model at least twice independently (basin shape requires nonlinearity).
    @irinarishRT @andrewjkoh: We develop a mechanism design framework for AI alignment and control: https://arxiv.org/abs/2609.01595 It’s largely conceptual but…
    @ghadfieldRT @knowledgeprob: The Bergemann, @andrewjkoh, Morris paper, alongside convos w/@alexolegimas @sebkrier @GordonBrianR this week, prompted m…

    10 Sources

    @andrewjkohWe develop a mechanism design framework for AI alignment and control: https://arxiv.org/abs/2609.01595 It’s largely conceptual but we offer stylized applications to failure modes (sandbagging, alignment faking), safety practice (scalable oversight, peer prediction), and a way to think about the value of alignment, interpretability, capability, and control.
    @alexolegimasFrom the newest member of our Economics team at @GoogleDeepMind: Mechanism design will be a key tool for alignment and safety with multi-agent systems. Andrew, Dirk, and Stephen present a rigorous framework for how to think about safety practices from a mechanism design perspective.
    @wsisaacRT @alexolegimas: From the newest member of our Economics team at @GoogleDeepMind: Mechanism design will be a key tool for alignment and sa…
    @malleshpaiReally nice paper!! There's a lovely line of work forming, this paper included, on applying mechanism design to AI alignment. AIs are autonomous black boxes that may well end up with preferences of their own. Alignment research mostly asks: how do we shape those preferences? But as models get more capable, it's getting harder and harder to directly shape those preferences and/or have confidence that we are shaping them as desired. Mechanism design asks the complementary question economists have asked about humans for decades: taking preferences as unknown and possibly bad, how do we design the rules (evals, permissions, rewards) so we get good outcomes anyway? We should be doing much more of the second.
    @sebkrierRT @andrewjkoh: We develop a mechanism design framework for AI alignment and control: https://arxiv.org/abs/2609.01595 It’s largely conceptual but…
    @ryan_t_lowedang, didn't realize that @andrewjkoh had joined DeepMind, great get
    @davidadThis is really cool and important work. Conjecture: this mechanism and the self-DPO I’ve been on about are instances of a common principle—to get a mind to fall into a robust basin, an RL scorer must use the model at least twice independently (basin shape requires nonlinearity).
    @irinarishRT @andrewjkoh: We develop a mechanism design framework for AI alignment and control: https://arxiv.org/abs/2609.01595 It’s largely conceptual but…
    @ghadfieldRT @knowledgeprob: The Bergemann, @andrewjkoh, Morris paper, alongside convos w/@alexolegimas @sebkrier @GordonBrianR this week, prompted m…