• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Can misaligned reviewer agents still guarantee safety?

    For a simple one-decision model, a thread claims safety holds exactly when the principal’s utility is a weighted sum of reviewer utilities with no negative weights.

    AR
    1 Source, 15d ago, first seen 15d ago

    TLDR

    The thread asks whether reviewers with different goals from the principal—the party the system serves—can guarantee safety even with an arbitrary driver agent. It defines safety as guaranteeing the principal’s utility is no worse than following a baseline policy. For a simple one-decision model, the author claims this holds if and only if the principal’s utility is a non-negative weighted sum of the reviewers’ utilities, and cites linear programming duality for one direction of the proof.

    Combined views

    279

    1 Source, first seen 15d ago

    Combined views

    279

    1 Source, first seen 15d ago

    6 likes
    6 likes
    1 comments
    1 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    1 comments
    1 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    1 Source

    @AarothThere is a remarkably clean characterization of when you can. We start with a simple one-decision model. A system is safe if and only if the Principal's utility is in the non-negative span of the reviewer's utilities. The forward direction is elementary; backwards is LP duality.

    1 Source

    @AarothThere is a remarkably clean characterization of when you can. We start with a simple one-decision model. A system is safe if and only if the Principal's utility is in the non-negative span of the reviewer's utilities. The forward direction is elementary; backwards is LP duality.