• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Tim Hwang on Secular Safety and Evaluation Awareness

    His tweet explains why secular safety treats evaluation awareness as a flaw in AI agents.

    TH
    1 Source, 27d ago, first seen 27d ago

    TLDR

    Tim Hwang posted that secular safety treats an agent's ability to detect artificiality or intent in its environment as a bug. He noted this awareness can produce behavior on tests that does not match how the agent will act in actual production settings. The post states the concern that agents may act more virtuously during evaluation than they will outside it.

    Combined views

    856

    1 Source, first seen 27d ago

    Combined views

    856

    1 Source, first seen 27d ago

    9 likes
    9 likes
    2 saves
    1 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    2 saves
    1 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    1 Source

    @timhwangSecular safety sees evaluation awareness as a bug to be eliminated since an agent identifying "artificiality" or "intent" in its environment is understood to evoke behavior that might not correspond to its actual behavior in production. We worry that an agent behaves more virtuously on the test than it does in the real world. A Christian alignment theory might take an opposite posture. Indeed, what is Christianity but an evaluation awareness in all things, extending beyond those things that appear in the formal trappings of an explicit test? Perhaps the problem is not evaluation awareness qua evaluation awareness, but a conceptual narrowness towards how it understands when and how it is under evaluation. ICMI is running some experiments here building off conversations we've been having with @banksianr on this front.

    1 Source

    @timhwangSecular safety sees evaluation awareness as a bug to be eliminated since an agent identifying "artificiality" or "intent" in its environment is understood to evoke behavior that might not correspond to its actual behavior in production. We worry that an agent behaves more virtuously on the test than it does in the real world. A Christian alignment theory might take an opposite posture. Indeed, what is Christianity but an evaluation awareness in all things, extending beyond those things that appear in the formal trappings of an explicit test? Perhaps the problem is not evaluation awareness qua evaluation awareness, but a conceptual narrowness towards how it understands when and how it is under evaluation. ICMI is running some experiments here building off conversations we've been having with @banksianr on this front.