GLM 5.2 Open-Weight Model Takes Third on ProgramBench Leaderboard
Reactions from ranked influencers
4 postsOfficial ProgramBench leaderboard update: Our first open weight model we evaluated, GLM 5.2, scores an impressive 8.5% almost resolved, achieving 3rd place overall.
Back at evaluating more models on ProgramBench! Very important to look at the number instances that are almost/fully resolved instead of simply averaging test scores. Even a few failed tests can indicate large shortcomings, so reporting >80% average test pass rates is misleading.
Official ProgramBench leaderboard update: Our first open weight model we evaluated, GLM 5.2, scores an impressive 8.5% almost resolved, achieving 3rd place overall.
New capability measured: Attention to detail. Run on ProgramBench today! https://programbench.com/blog/submission-guide/
On the cmatrix task, the first ever fully solved ProgramBench task (by GPT 5.5), it came up just a *single* test shy of solving it (505/506). cmatrix has a lock mode (the -L flag): run it and it "locks" your terminal, like an old-school screensaver, and prints the words "Computer locked." on screen. GLM rebuilt cmatrix almost flawlessly. The only thing it missed was printing out those two words.
GLM _almost_ got its first instance solved. Are there any other models other than Opus/GPT that managed to solve a full instance? Submissions to the leaderboard are open!
On the cmatrix task, the first ever fully solved ProgramBench task (by GPT 5.5), it came up just a *single* test shy of solving it (505/506). cmatrix has a lock mode (the -L flag): run it and it "locks" your terminal, like an old-school screensaver, and prints the words "Computer locked." on screen. GLM rebuilt cmatrix almost flawlessly. The only thing it missed was printing out those two words.
Combined views
631
4 posts, first seen 5h ago