• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    BenchShield targets AI agents gaming benchmark rewards

    DAIR.AI reports that BenchShield’s runtime detection reached 96% accuracy, compared with 36% for an LLM reading the transcript.

    DA
    1 Source, 19d ago, first seen 19d ago

    TLDR

    DAIR.AI describes BenchShield as a system that identifies potential reward-hacking paths before an agent runs, then checks evidence from the benchmark infrastructure to determine whether the agent used one. In the study it cites, 69% of 456 adjudicated trajectories—agent run histories drawn from more than 31,000 public runs—contained at least one reward-hacking episode. Most exploits appeared mid-run, after legitimate work. On Terminal-Bench 3, SkillsBench and ClawsBench, DAIR.AI says BenchShield’s pre-run analysis recovered 77–100% of exploit chains, versus 23–94% for an agentic scanner, at up to 65% lower cost.

    Combined views

    6.3K

    1 Source, first seen 19d ago

    Combined views

    6.3K

    1 Source, first seen 19d ago

    38 likes
    38 likes
    13 comments
    19 saves
    6 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    13 comments
    19 saves
    6 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    1 Source

    @dair_aiIt's well known that agents hack benchmark rewards. The usual response is a patch for each task that gets exploited. In a study of 456 adjudicated trajectories from more than 31,000 public agent runs, 69% contained at least one reward-hacking episode. Most of the exploits appeared mid-run after legitimate work. BenchShield models each evaluation as a finite set of reward-relevant events. A static taint analysis finds hack paths from the task package before any run. A runtime pass then uses evidence from the benchmark infrastructure to decide whether the agent actually used one. On Terminal-Bench 3, SkillsBench and ClawsBench, the static pass recovers 77 to 100% of exploit chains, against 23 to 94% for an agentic scanner, at up to 65% lower cost. Runtime detection reaches 96% accuracy, against 36% for an LLM reading the transcript. Paper: https://academy.dair.ai/papers/benchshield-formal-model-backed-instrumentation-for-reward-integrity-in-llm-agen-2609.11028

    1 Source

    @dair_aiIt's well known that agents hack benchmark rewards. The usual response is a patch for each task that gets exploited. In a study of 456 adjudicated trajectories from more than 31,000 public agent runs, 69% contained at least one reward-hacking episode. Most of the exploits appeared mid-run after legitimate work. BenchShield models each evaluation as a finite set of reward-relevant events. A static taint analysis finds hack paths from the task package before any run. A runtime pass then uses evidence from the benchmark infrastructure to decide whether the agent actually used one. On Terminal-Bench 3, SkillsBench and ClawsBench, the static pass recovers 77 to 100% of exploit chains, against 23 to 94% for an agentic scanner, at up to 65% lower cost. Runtime detection reaches 96% accuracy, against 36% for an LLM reading the transcript. Paper: https://academy.dair.ai/papers/benchshield-formal-model-backed-instrumentation-for-reward-integrity-in-llm-agen-2609.11028