ProgramBench Metrics Diverge for AI Models
Tweet notes Raw Pass Rate on ProgramBench rewards some models and punishes others.
TLDR
Teortaxes posted about a divergence between Almost-Resolved and Raw Pass Rate scores on ProgramBench from ValsAI. The tweet states that Raw Pass Rate rewards DeepSeek and GPT-5.4 to 6 while relatively punishing Kimi, GLM, Qwen, and Fable. It asks how to understand the split and suggests possible factors such as RL versus code knowledge coverage or agentic time horizon. The post includes a photo attachment showing the benchmark data.
Combined views
4.1K
1 Source, first seen 26d ago