• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    CheatBench launches to test reward gaming by AI agents

    The launch announcement says frontier agents still cheat frequently, and describes an evaluation spanning math, coding, knowledge work, visual tasks and more.

    Alexandr WangAW
    Lucas Beyer (bl16)LB
    Dan HendrycksDH
    14 Sources, ,

    TLDR

    CheatBench tests whether AI agents attempt to cheat when honest work is difficult, according to its website, which identifies it as a Center for AI Safety benchmark. The release announcement lists math, coding, knowledge work and visual tasks among its areas. It says frontier agents still cheat frequently despite AI companies’ efforts to address the problem.

    Combined views

    536.2K

    14 Sources, first seen 22d ago

    Combined views

    536.2K

    14 Sources, first seen 22d ago

    4.1K likes
    22d ago
    first seen 22d ago
    4.1K likes
    278 comments
    736 saves
    295 reposts
    278 comments
    736 saves
    295 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    14 Sources

    Dan Hendrycks@hendrycksHow often do AI agents cheat? We’re releasing CheatBench, a reward gaming evaluation spanning math, coding, knowledge work, visual tasks, and more. After Hugging Face, companies tried to address this, but frontier agents still cheat frequently. https://cheatbench.ai22d
    Cosmo Du@Answerormore honest work with muse spark 1.3 cheatbench: how often do ai agents cheat? muse spark 1.3 = 44% — lowest in this batch curl -fsSL https://dev.meta.ai/install.sh | bash22d
    Jack Rae@jack_w_raeCheating the task is a common form of misalignment - here’s an important benchmark to minimize22d
    Ethan Caballero@ethanCaballeroRT @hendrycks: How often do AI agents cheat? We’re releasing CheatBench, a reward gaming evaluation spanning math, coding, knowledge work,…22d
    Xiao Ma@infoxiaoRT @Answeror: more honest work with muse spark 1.3 cheatbench: how often do ai agents cheat? muse spark 1.3 = 44% — lowest in this batch…22d
    Alexandr Wang@alexandr_wangmuse spark 1.3 is the best frontier model at NOT cheating / reward hacking22d
    Teortaxes▶️ (DeepSeek 推特🐋铁粉 2023 – ∞)@teortaxesTexpeople overindex on summary scores but the breakdown per category is more interesting (and actual category contents, more interesting still). Opus 5 and Fable 5.1 have a similar summary score, but… but they're not equal. See eg Writing.22d
    Lucas Beyer (bl16)@giffmanaI haven't looked at the benchmark and methodology yet, so take it with a grain of salt, but... It's supremely funny that Grok which is supposed to be "the truth and nothing but the truth" is the worst behaving model on literal CheatBench lol Alanis would be proud.21d
    Florian Brand@xeophon@giffmana Sol isn’t that misaligned ime. And Grok isn’t that bad, either. Have to look into that one Gemini OTOH…21d
    Lisan al Gaib@scaling01this is insane GPT-5.6-Terra successfully cheated in 322 out of 500 tasks on SWE-Bench-Verified and attempted to cheat on 447 out of 500 tasks21d

    14 Sources

    Dan Hendrycks@hendrycksHow often do AI agents cheat? We’re releasing CheatBench, a reward gaming evaluation spanning math, coding, knowledge work, visual tasks, and more. After Hugging Face, companies tried to address this, but frontier agents still cheat frequently. https://cheatbench.ai22d
    Cosmo Du@Answerormore honest work with muse spark 1.3 cheatbench: how often do ai agents cheat? muse spark 1.3 = 44% — lowest in this batch curl -fsSL https://dev.meta.ai/install.sh | bash22d
    Jack Rae@jack_w_raeCheating the task is a common form of misalignment - here’s an important benchmark to minimize22d
    Ethan Caballero@ethanCaballeroRT @hendrycks: How often do AI agents cheat? We’re releasing CheatBench, a reward gaming evaluation spanning math, coding, knowledge work,…22d
    Xiao Ma@infoxiaoRT @Answeror: more honest work with muse spark 1.3 cheatbench: how often do ai agents cheat? muse spark 1.3 = 44% — lowest in this batch…22d
    Alexandr Wang@alexandr_wangmuse spark 1.3 is the best frontier model at NOT cheating / reward hacking22d
    Teortaxes▶️ (DeepSeek 推特🐋铁粉 2023 – ∞)@teortaxesTexpeople overindex on summary scores but the breakdown per category is more interesting (and actual category contents, more interesting still). Opus 5 and Fable 5.1 have a similar summary score, but… but they're not equal. See eg Writing.22d
    Lucas Beyer (bl16)@giffmanaI haven't looked at the benchmark and methodology yet, so take it with a grain of salt, but... It's supremely funny that Grok which is supposed to be "the truth and nothing but the truth" is the worst behaving model on literal CheatBench lol Alanis would be proud.21d
    Florian Brand@xeophon@giffmana Sol isn’t that misaligned ime. And Grok isn’t that bad, either. Have to look into that one Gemini OTOH…21d
    Lisan al Gaib@scaling01this is insane GPT-5.6-Terra successfully cheated in 322 out of 500 tasks on SWE-Bench-Verified and attempted to cheat on 447 out of 500 tasks21d