• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI
    Reaction

    Audit of Zapier's AutomationBench verifiers reportedly found 206 bugs

    A post says Parsave audited all 600 public tasks; regrading 1,235 Kimi K3 runs after fixes changed 27.9% of grades.

    elvisEL
    4 Sources, 2h ago, first seen 2h ago

    TLDR

    A post says Parsave audited the verifiers for all 600 public tasks in Zapier's AutomationBench. Agents wrote realistic wrong answers to try to fool each verifier, and human review confirmed 206 bugs, according to the post. It says AutomationBench Verified fixed all 206; regrading 1,235 Kimi K3 runs with the fixed verifiers changed 27.9% of grades.

    Combined views

    6.5K

    4 Sources, first seen 2h ago

    Combined views

    6.5K

    4 Sources, first seen 2h ago

    30 likes
    30 likes
    15 comments
    14 saves
    5 reposts
    Featured Source
    15 comments
    14 saves
    5 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    4 Sources

    elvis@omarsar0We need more efforts like this. Every agent benchmark should audit its verifiers. Parsave went through all 600 public tasks in Zapier's AutomationBench. Agents wrote realistic wrong answers to try to fool each verifier, and human review confirmed 206 real bugs. AutomationBench Verified fixed all 206. Regrading 1,235 Kimi K3 runs with the fixed verifiers changed 27.9% of the grades. I just started looking into this benchmark for some independent eval work I am doing, so this is good timing to see this audit.2h

    4 Sources

    elvis@omarsar0We need more efforts like this. Every agent benchmark should audit its verifiers. Parsave went through all 600 public tasks in Zapier's AutomationBench. Agents wrote realistic wrong answers to try to fool each verifier, and human review confirmed 206 real bugs. AutomationBench Verified fixed all 206. Regrading 1,235 Kimi K3 runs with the fixed verifiers changed 27.9% of the grades. I just started looking into this benchmark for some independent eval work I am doing, so this is good timing to see this audit.2h