• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Models Show Pass@3 Drops on Benchmark Variants

    Fidian tested top models on a variant of Terminal-Bench using the same 89 tasks.

    AB
    1 Source, 32d ago, first seen 32d ago

    TLDR

    Ahmad Beirami, a research engineer focused on RL and LLM post-training, posted that some models exhibit significant pass@3 drops when tasks are replaced with variants. The claim references tests by Fidian that ran the top 20 models from Artificial Analysis's Terminal-Bench 2.1 leaderboard on both the original benchmark and TB-fn. TB-fn is Fidian's variant built from the same 89 tasks as TB-2.1. The post calls the drops suspicious.

    Combined views

    1 Source, first seen 32d ago

    Combined views

    1 Source, first seen 32d ago

    27 reposts
    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    27 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    1 Source

    @abeiramiRT @abeirami: Suspicious! We had found that some models exhibit significant pass@3 drops when you replace the tasks with variants that hav…

    1 Source

    @abeiramiRT @abeirami: Suspicious! We had found that some models exhibit significant pass@3 drops when you replace the tasks with variants that hav…