MiniCPM5-2B scores 831 on work-task benchmark, Artificial Analysis reports
Artificial Analysis puts the model below GDPval-AA v2's human baseline of 1,000, but ahead of Ling 3.0 Tiny and Granite 4.2 8B on the same test.
TLDR
Artificial Analysis reports an Elo score of 831 for MiniCPM5-2B on GDPval-AA v2, which tests real-world work tasks against a human baseline of 1,000. Its reported score exceeds Ling 3.0 Tiny's 718 and Granite 4.2 8B's 647.
On AA-Briefcase, its evaluation of agentic knowledge work, Artificial Analysis says MiniCPM5-2B ranks second in the comparison set at 438, behind Ling 3.0 Tiny (485) and ahead of Granite 4.2 8B (324).
Combined views
10.3K
3 Sources, first seen 23d ago