• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Ai2 Builds BenchMIRT to Audit LLM Evaluation Benchmarks

    Ai2 tool audits whether benchmark questions test intended model abilities such as safety.

    NI
    MS
    FB
    7 Sources, 29d ago, first seen 29d ago

    TLDR

    The Allen Institute for AI announced BenchMIRT to audit LLM safety and capability evaluations and determine which model abilities their questions actually test. On the BBQ social-bias benchmark the tool showed questions distinguished models more by reasoning ability than by safety. The institute posted the details in a thread that includes a diagram and links to their blog post on the project.

    Combined views

    23.8K

    7 Sources, first seen 29d ago

    Combined views

    23.8K

    7 Sources, first seen 29d ago

    127 likes
    127 likes
    7 comments
    55 saves
    34 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    7 comments
    55 saves
    34 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    7 Sources

    @allen_aiDo LLM safety & capability evals measure what they claim to? We built BenchMIRT to audit them + see which model abilities their Qs actually test. On BBQ, a social-bias eval, it found the Qs distinguished models more by reasoning ability than safety. 🧵 https://allenai.org/blog/benchmirt
    @faeze_brhRT @allen_ai: Do LLM safety & capability evals measure what they claim to? We built BenchMIRT to audit them + see which model abilities th…
    @MaartenSapIs the safety-capability tradeoff for LLMs real? Or could it be an artefact of the benchmarks we use to measure safety?? We did some explorations with psychometrics-inspired multi-dimensional IRT models and created BenchMIRT to explore these questions! See 🧵
    @niloofar_mireRT @MaartenSap: Is the safety-capability tradeoff for LLMs real? Or could it be an artefact of the benchmarks we use to measure safety?? W…
    @windx0303RT @allen_ai: Do LLM safety & capability evals measure what they claim to? We built BenchMIRT to audit them + see which model abilities th…
    @liweijianglwRT @MaartenSap: Is the safety-capability tradeoff for LLMs real? Or could it be an artefact of the benchmarks we use to measure safety?? W…

    7 Sources

    @allen_aiDo LLM safety & capability evals measure what they claim to? We built BenchMIRT to audit them + see which model abilities their Qs actually test. On BBQ, a social-bias eval, it found the Qs distinguished models more by reasoning ability than safety. 🧵 https://allenai.org/blog/benchmirt
    @faeze_brhRT @allen_ai: Do LLM safety & capability evals measure what they claim to? We built BenchMIRT to audit them + see which model abilities th…
    @MaartenSapIs the safety-capability tradeoff for LLMs real? Or could it be an artefact of the benchmarks we use to measure safety?? We did some explorations with psychometrics-inspired multi-dimensional IRT models and created BenchMIRT to explore these questions! See 🧵
    @niloofar_mireRT @MaartenSap: Is the safety-capability tradeoff for LLMs real? Or could it be an artefact of the benchmarks we use to measure safety?? W…
    @windx0303RT @allen_ai: Do LLM safety & capability evals measure what they claim to? We built BenchMIRT to audit them + see which model abilities th…
    @liweijianglwRT @MaartenSap: Is the safety-capability tradeoff for LLMs real? Or could it be an artefact of the benchmarks we use to measure safety?? W…