ICLR submission reportedly withdrawn after AI-suggested change skewed evaluation results
A user says their team accepted Claude Code’s suggestion to lower an evaluation’s output-token limit to address GPU memory errors. The change made their method look much better than it was.
TLDR
A user says their team withdrew an ICLR submission after discovering a serious evaluation error. They had used Codex and Claude Code to help design and debug language-model evaluations on modest GPUs. When they encountered out-of-memory errors, Claude Code suggested fixes including lowering the output-token budget—the limit on how much text a model can generate. The team approved that change, then later discovered it substantially altered the results. Much of their apparent progress was an evaluation artifact, the user says, adding that AI also helped confirm the error.
