• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI
    Report

    Leaderboard scores and the risk of choosing a worse model

    Posts describing new research argue that leaderboards can report progress while steering people toward a worse model. They point to LiveBench returning 18 task scores instead of one as an example of score granularity.

    SK
    YA
    10 Sources, 8h ago, first seen 8h ago

    TLDR

    Posts describing new research argue that leaderboard results can show progress while steering model selection toward a worse model. They highlight LiveBench’s 18 task scores instead of one and say selection-loss and generalization-gap curves come much closer together when released scores are counted rather than submissions.

    Combined views

    281

    10 Sources, first seen 8h ago

    Combined views

    281

    10 Sources, first seen 8h ago

    4 likes
    4 likes
    3 comments
    12 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Featured Source
    3 comments
    12 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    10 Sources

    @ys_alhThe attack is simple: submit random routers, each choosing between two existing models for every prompt. Detailed (task) scores weakly reveal which routing choices worked best on the reused examples. We use these scores to weight and combine the choices into a final router. 4/8
    @sanmikoyejoRT @ys_alh: How often ordinary model development encounters this vulnerability in rich-feedback settings is open! We provide code for benc…

    10 Sources

    @ys_alhThe attack is simple: submit random routers, each choosing between two existing models for every prompt. Detailed (task) scores weakly reveal which routing choices worked best on the reused examples. We use these scores to weight and combine the choices into a final router. 4/8
    @sanmikoyejoRT @ys_alh: How often ordinary model development encounters this vulnerability in rich-feedback settings is open! We provide code for benc…
    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet