• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    “Coalitional alignment” and reward-model safety/completeness tradeoffs

    A research thread describes a condition weaker than requiring reviewer agents to match your utility exactly—and cautions that it is not guaranteed to hold.

    AR
    1 Source, 15d ago, first seen 15d ago

    TLDR

    The thread’s author calls the condition “coalitional alignment,” describing it as substantially weaker than individual alignment, where reviewer agents have exactly your utility. Even that weaker condition is not guaranteed to hold. The author reports some preliminary evidence of non-trivial safety/completeness tradeoffs among real reward models, attributing those effects to coalitional rather than individual alignment.

    Combined views

    355

    1 Source, first seen 15d ago

    Combined views

    355

    1 Source, first seen 15d ago

    3 likes
    3 likes
    1 comments
    1 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    1 comments
    1 reposts
    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    1 Source

    @AarothBut despite being weaker it is not guaranteed to be satisfied. Still, we find some preliminary evidence of non-trivial safety/completeness tradeoffs among real reward models, and these effects come from coallitional rather than individual alignment.

    1 Source

    @AarothBut despite being weaker it is not guaranteed to be satisfied. Still, we find some preliminary evidence of non-trivial safety/completeness tradeoffs among real reward models, and these effects come from coallitional rather than individual alignment.