Announcement
Decision grader is reportedly ~32× cheaper than an LLM judge in a code QA test
A user describing Applied Compute’s test says it was ~8× faster, with 94% agreement across 407 criteria.
TLDR
A user describing Applied Compute’s AC2 code QA benchmark says a decision grader was ~32× cheaper and ~8× faster than an LLM judge on the same answers across 100 sample tasks (407 criteria), with 94% agreement. It returned yes/no probabilities for every criterion in one call. The user says confident results matched the LLM judge 99% of the time, while low-confidence results helped identify ambiguous rubric items.
Combined views
9.5K
2 Sources, first seen ago
