• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    More benchmarks and more voices in AI testing

    A post makes the case for “eval pluralism”: more benchmarks, input from clinicians, engineers, scientists and everyday users, and open infrastructure for evaluating AI.

    OP
    AR
    VS
    5 Sources, ,

    TLDR

    A post argues that a single benchmark measures only a slice of what an AI system can do. It calls for broader testing across environments, outputs and levels of autonomy—from simple prompts to full worlds, and single turns to continually improving agents. It also argues that subject-matter experts and everyday users, not just a small group of AI researchers, should help define what “good” looks like, supported by shared, open evaluation infrastructure.

    Combined views

    14.4K

    5 Sources, first seen 23d ago

    Combined views

    14.4K

    5 Sources, first seen 23d ago

    114 likes
    23d ago
    first seen 23d ago
    114 likes
    14 comments
    25 saves
    7 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    14 comments
    25 saves
    7 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    5 Sources

    @ajratnerFor rate/robustness: a great example of progress here is the idea of "continuous benchmarks" from @ryan_marten
    @OfirPressi remember 3.5 years ago HumanEval was all the rage. i had already created one benchmark (Bamboogle) and was thinking about the next and thought "there's *no way* anyone will be able to think of a bench that's better or more elegant or more challenging than HumanEval"
    @vincentsunnchenstrongly agree that we need "a much broader and more ingenious stable of evaluations." getting there will require eval pluralism: more benchmarks, more evaluators (including subject-matter experts), and open eval infrastructure - more benchmarks: a single benchmark measures a slice of what these systems can do, offering an empirical tool to understand capabilities, risks, and failure modes. earlier this year we shared our view of the axes that need coverage: environment complexity (prompt/response → full worlds), output complexity (fixed-schema labels → nuanced, subjective artifacts), and autonomy horizon (single turns → continually improving agents): https://x.com/vincentsunnchen/status/2021659820240384141. @ajratner shared his mental model on 'smoothing' eval coverage: - more evaluators: critically, measurement shouldn't sit only with a small group of AI researchers. we are encouraged by the increasing number of independent researchers working on evaluation / data, and believe even more subject-matter expertise will be needed to close the evaluation gap. people with the most to gain (and lose!) from these systems should be the ones defining what "good" looks like: clinicians, engineers, scientists, and everyday users - open eval infrastructure: the above requires a shared infrastructure built with people and technology. we love @ryan_marten's framing of "benchmarks are software" (https://x.com/ryan_marten/status/2080321791361527843) and see room for continued investment in continuous QC, grading design, and reward-hacking monitoring, esp as models show increasing 'situational awareness' (h/t @henryehrenberg: https://senior-swe-bench.snorkel.ai/blog/2026-07-07-fable-5-top-spot#claude-fable-5-demonstrates-far-more-situational-awareness) Open Benchmarks Grants (https://benchmarks.snorkel.ai) is just the start of our push to support all of the above - and we have a lot more coming soon. please reach out if you are building here - we would love to work with you!

    5 Sources

    @ajratnerFor rate/robustness: a great example of progress here is the idea of "continuous benchmarks" from @ryan_marten
    @OfirPressi remember 3.5 years ago HumanEval was all the rage. i had already created one benchmark (Bamboogle) and was thinking about the next and thought "there's *no way* anyone will be able to think of a bench that's better or more elegant or more challenging than HumanEval"
    @vincentsunnchenstrongly agree that we need "a much broader and more ingenious stable of evaluations." getting there will require eval pluralism: more benchmarks, more evaluators (including subject-matter experts), and open eval infrastructure - more benchmarks: a single benchmark measures a slice of what these systems can do, offering an empirical tool to understand capabilities, risks, and failure modes. earlier this year we shared our view of the axes that need coverage: environment complexity (prompt/response → full worlds), output complexity (fixed-schema labels → nuanced, subjective artifacts), and autonomy horizon (single turns → continually improving agents): https://x.com/vincentsunnchen/status/2021659820240384141. @ajratner shared his mental model on 'smoothing' eval coverage: - more evaluators: critically, measurement shouldn't sit only with a small group of AI researchers. we are encouraged by the increasing number of independent researchers working on evaluation / data, and believe even more subject-matter expertise will be needed to close the evaluation gap. people with the most to gain (and lose!) from these systems should be the ones defining what "good" looks like: clinicians, engineers, scientists, and everyday users - open eval infrastructure: the above requires a shared infrastructure built with people and technology. we love @ryan_marten's framing of "benchmarks are software" (https://x.com/ryan_marten/status/2080321791361527843) and see room for continued investment in continuous QC, grading design, and reward-hacking monitoring, esp as models show increasing 'situational awareness' (h/t @henryehrenberg: https://senior-swe-bench.snorkel.ai/blog/2026-07-07-fable-5-top-spot#claude-fable-5-demonstrates-far-more-situational-awareness) Open Benchmarks Grants (https://benchmarks.snorkel.ai) is just the start of our push to support all of the above - and we have a lot more coming soon. please reach out if you are building here - we would love to work with you!