Fixed Eval Boosts Open Model Over Later Closed Releases
Research engineer reports performance change after correcting benchmark errors.
Florian Brand noted that fixing an evaluation produced a large performance increase for an open model released before the benchmark. A closed model released afterward remained unchanged. Cody Blakeney audited MBPP and HumanEval variants and corrected issues with prompts, test cases, and formatting. He used an NAC tool to run concurrent sessions and plans to open-source the fixed evaluations along with further details in a blog post.
Combined views
9.9K
7 posts, first seen 4d ago