The case for “anti-slop” AI coding evaluations
A user argues that unwanted AI-generated tests can fail code review even when a model scores well on a benchmark—and that human review and cleanup should count as failure.
TLDR
A user calls for “anti-slop” coding evaluations that account for whether generated code would pass review. They describe Claude adding tests they believe no engineer would keep, contrasting that with benchmark performance. Their proposed standard: if a human has to review and clean up the output, it should receive a failing score.
The case for “anti-slop” AI coding evaluations
A user argues that unwanted AI-generated tests can fail code review even when a model scores well on a benchmark—and that human review and cleanup should count as failure.
TLDR
A user calls for “anti-slop” coding evaluations that account for whether generated code would pass review. They describe Claude adding tests they believe no engineer would keep, contrasting that with benchmark performance. Their proposed standard: if a human has to review and clean up the output, it should receive a failing score.