• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI
    Reaction

    Frontier AI companies reportedly commit to employee-like access for outside evaluators

    Apollo Research says access to training, evaluation and deployment could matter, but its impact depends heavily on implementation.

    AL
    1 Source, 15h ago, first seen 15h ago

    TLDR

    Apollo Research says frontier AI companies have committed to giving outside evaluators employee-like access to training, evaluation and deployment. It shared principles for embedded evaluations, stressing that implementation will determine their impact. In a separate post, a commenter predicts evaluators may eventually shift from finding obvious issues to testing whether developers would have caught problems if they arose.

    Combined views

    1.6K

    1 Source, first seen 15h ago

    Combined views

    1.6K

    1 Source, first seen 15h ago

    22 likes
    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    22 likes
    13 saves
    3 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Featured Source
    13 saves
    3 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    1 Source

    @AlexMeinkeI think this framework for evaluating safety claims is actually really elegant. I expect right now embedded evaluators will be finding countless things that are directly on fire But once we stop finding obvious issues everywhere, the work of embedded evaluators should shift towards carefully red-teaming whether a developer *would* have caught problems *if* they occurred The three-step process we outline there directly incentivizes this hill-climbing on safety standards

    1 Source

    @AlexMeinkeI think this framework for evaluating safety claims is actually really elegant. I expect right now embedded evaluators will be finding countless things that are directly on fire But once we stop finding obvious issues everywhere, the work of embedded evaluators should shift towards carefully red-teaming whether a developer *would* have caught problems *if* they occurred The three-step process we outline there directly incentivizes this hill-climbing on safety standards