• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    A deeper bench for AI evaluation beyond METR

    The post says strict nondisclosure agreements keep many evaluators from discussing their work publicly, even as they contribute to AI system and model documentation.

    CM
    MB
    RO
    49 Sources, ,

    TLDR

    A user calls for more talent at existing AI evaluators, new organizations across areas of expertise, and stronger standards and frameworks for frontier AI auditing. The post argues that voluntary commitments are not enough: standards, regulation, and frameworks could give evaluators more consistent, durable operating conditions.

    Combined views

    447.4K

    49 Sources, first seen 23d ago

    Combined views

    447.4K

    49 Sources, first seen 23d ago

    5.5K likes
    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    23d ago
    first seen 23d ago
    5.5K likes
    379 comments
    1.1K saves
    477 reposts
    379 comments
    1.1K saves
    477 reposts

    Sentiment

    Positive48.1%51.9%Negative

    Summary

    Sentiment

    Positive48.1%51.9%Negative

    Positive accounts welcomed METR gaining access to Anthropic when paired with diverse independent evaluators to reduce single-worldview bias, while negative replies called METR a non-independent EA front with weak model tracking.

    Based on 296 sentiment-bearing replies from 241 accounts across 15 conversations.

    Summary

    Positive accounts welcomed METR gaining access to Anthropic when paired with diverse independent evaluators to reduce single-worldview bias, while negative replies called METR a non-independent EA front with weak model tracking.

    Based on 296 sentiment-bearing replies from 241 accounts across 15 conversations.

    49 Sources

    @ChrisPainterYupI’m obviously biased but I think the theory of change of “ensure that enough progress has been made on alignment*, monitor this across industry, and inform the public if it hasn’t” makes at least as much sense for many people to try as “align the models by directly working on alignment at a lab” And many people at labs don’t work on alignment! *as a specific case of the more narrow “check whether rogue AI is a big immediate risk factor” which is often more precisely the thing a lot of our outputs measure
    @parallx_aiIntroducing Parallax 🔍 We're a new UK/EU-based nonprofit research lab building scalable methods and infrastructure to audit the beliefs, goals, and plans that shape model behaviours in long-horizon agentic evaluations. https://parallx.ai — Thread 🧵 1/
    @boknilevRT @parallx_ai: Introducing Parallax 🔍 We're a new UK/EU-based nonprofit research lab building scalable methods and infrastructure to audi…
    @prpaskovAgree re: METR; but we also need a deeper bench! People may be surprised to learn that the evaluator ecosystem already runs wider than the current (well-deserved) focus on METR suggests. Many strong evaluation orgs operate under tight NDAs and can't speak publicly to their work, despite making core and continued evaluation contributions to system and model cards. Some of these orgs you may have never heard of! This is one example why we need more than voluntary commitments. Standards, regulation, and frameworks can create more consistent and durable operating conditions for all evaluators. We need: 1. More talent flowing into other existing evaluators 2. New orgs across domains of expertise for a more diversified ecosystem 3. Stronger standards and frameworks for frontier AI auditing
    @wsisaacRT @prpaskov: Agree re: METR; but we also need a deeper bench! People may be surprised to learn that the evaluator ecosystem already runs…
    @logangrahamLet a thousand METRs bloom. If you're a founder type and AI safety-curious, maybe you should start an independent auditor/evaluator. Happy to help w/ advice/connections/maybe $!
    @_sholtodouglas100%, this is going to be one of the most important roles in the world - the people will need to be of incredible integity and technical expertise, and be drawn from a wide enough set of backgrounds that all of society feels confidence in their judgements.
    @_arohan_@tszzl @_sholtodouglas I think this is good. But no one buys the independence of the independent evaluator idea just yet is what I understand reading the news.
    @anton_d_leichtin scaling this ecosystem, we need to make sure it draws on a wide range of funding and views on AI risk, anchored in common standards. we can't afford to squander this rare moment of trust by getting territorial or trying to keep 'outsiders' out. look at how we got to the current ecosystem: until very recently, there was only a small community interested in frontier risk evals. the current orgs in the space were simply the most far-sighted about where this was going. that means it's a pretty small selection to start with. but now that there's buy-in from the labs, that's the demand signal the ecosystem needs to mature beyond those that came early: it shows to talent and money that this is worth pursuing and worth funding. the independent evaluation ecosystem has rightly been given the benefit of the doubt; now, it's really important to remain above reproach and show the world that that was the right call.
    @nbaschezI am all for this, but there is a potentially big incentive issue that I haven’t seen anyone talking about yet If the labs select and pay the evaluators, they have a strong disincentive to piss off the labs or make them look bad It’s the same incentive problem that contributed to the GFC, when Moody’s, etc rated subprime mortgage backed securities as AAA.

    49 Sources

    @ChrisPainterYupI’m obviously biased but I think the theory of change of “ensure that enough progress has been made on alignment*, monitor this across industry, and inform the public if it hasn’t” makes at least as much sense for many people to try as “align the models by directly working on alignment at a lab” And many people at labs don’t work on alignment! *as a specific case of the more narrow “check whether rogue AI is a big immediate risk factor” which is often more precisely the thing a lot of our outputs measure
    @parallx_aiIntroducing Parallax 🔍 We're a new UK/EU-based nonprofit research lab building scalable methods and infrastructure to audit the beliefs, goals, and plans that shape model behaviours in long-horizon agentic evaluations. https://parallx.ai — Thread 🧵 1/
    @boknilevRT @parallx_ai: Introducing Parallax 🔍 We're a new UK/EU-based nonprofit research lab building scalable methods and infrastructure to audi…
    @prpaskovAgree re: METR; but we also need a deeper bench! People may be surprised to learn that the evaluator ecosystem already runs wider than the current (well-deserved) focus on METR suggests. Many strong evaluation orgs operate under tight NDAs and can't speak publicly to their work, despite making core and continued evaluation contributions to system and model cards. Some of these orgs you may have never heard of! This is one example why we need more than voluntary commitments. Standards, regulation, and frameworks can create more consistent and durable operating conditions for all evaluators. We need: 1. More talent flowing into other existing evaluators 2. New orgs across domains of expertise for a more diversified ecosystem 3. Stronger standards and frameworks for frontier AI auditing
    @wsisaacRT @prpaskov: Agree re: METR; but we also need a deeper bench! People may be surprised to learn that the evaluator ecosystem already runs…
    @logangrahamLet a thousand METRs bloom. If you're a founder type and AI safety-curious, maybe you should start an independent auditor/evaluator. Happy to help w/ advice/connections/maybe $!
    @_sholtodouglas100%, this is going to be one of the most important roles in the world - the people will need to be of incredible integity and technical expertise, and be drawn from a wide enough set of backgrounds that all of society feels confidence in their judgements.
    @_arohan_@tszzl @_sholtodouglas I think this is good. But no one buys the independence of the independent evaluator idea just yet is what I understand reading the news.
    @anton_d_leichtin scaling this ecosystem, we need to make sure it draws on a wide range of funding and views on AI risk, anchored in common standards. we can't afford to squander this rare moment of trust by getting territorial or trying to keep 'outsiders' out. look at how we got to the current ecosystem: until very recently, there was only a small community interested in frontier risk evals. the current orgs in the space were simply the most far-sighted about where this was going. that means it's a pretty small selection to start with. but now that there's buy-in from the labs, that's the demand signal the ecosystem needs to mature beyond those that came early: it shows to talent and money that this is worth pursuing and worth funding. the independent evaluation ecosystem has rightly been given the benefit of the doubt; now, it's really important to remain above reproach and show the world that that was the right call.
    @nbaschezI am all for this, but there is a potentially big incentive issue that I haven’t seen anyone talking about yet If the labs select and pay the evaluators, they have a strong disincentive to piss off the labs or make them look bad It’s the same incentive problem that contributed to the GFC, when Moody’s, etc rated subprime mortgage backed securities as AAA.