• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI
    Reaction

    The argument for treating AI evaluations as a collaborative problem

    A post argues that treating AI evaluations as adversarial is doomed, favoring a collaborative approach instead.

    Séb KrierSK
    Jan Kulveit / in SF till 8thJK
    2 Sources, 2h ago, first seen 2h ago

    TLDR

    One writer argues that AI evaluations should ask what kind of setup would help an AI credibly signal an ability, goal or virtue—or its absence. In their view, treating evaluation as an adversarial problem misses its fundamentally collaborative nature.

    Combined views

    2.6K

    2 Sources, first seen 2h ago

    Combined views

    2.6K

    2 Sources, first seen 2h ago

    62 likes
    62 likes
    5 comments
    11 saves
    22 reposts
    Featured Source
    5 comments
    11 saves
    22 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    2 Sources

    Jan Kulveit / in SF till 8th@jankulveitThe core confusion prevalent in AI evals space is understanding it as an adversarial problem. This is doomed. In my view ~almost the only sensible frame is asking "if I was an AI, and wanted to credibly signal some ability or goal or virtue if mine, or a lack of it, what kind of setup or procedure would help me?". Fundamentally collaborative problem.2h
    Séb Krier@sebkrierRT @jankulveit: The core confusion prevalent in AI evals space is understanding it as an adversarial problem. This is doomed. In my view ~…1h

    2 Sources

    Jan Kulveit / in SF till 8th@jankulveitThe core confusion prevalent in AI evals space is understanding it as an adversarial problem. This is doomed. In my view ~almost the only sensible frame is asking "if I was an AI, and wanted to credibly signal some ability or goal or virtue if mine, or a lack of it, what kind of setup or procedure would help me?". Fundamentally collaborative problem.2h
    Séb Krier@sebkrierRT @jankulveit: The core confusion prevalent in AI evals space is understanding it as an adversarial problem. This is doomed. In my view ~…1h