• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    MiniCPM5-2B scores 831 on work-task benchmark, Artificial Analysis reports

    Artificial Analysis puts the model below GDPval-AA v2's human baseline of 1,000, but ahead of Ling 3.0 Tiny and Granite 4.2 8B on the same test.

    T(
    AA
    3 Sources, ,

    TLDR

    Artificial Analysis reports an Elo score of 831 for MiniCPM5-2B on GDPval-AA v2, which tests real-world work tasks against a human baseline of 1,000. Its reported score exceeds Ling 3.0 Tiny's 718 and Granite 4.2 8B's 647.

    On AA-Briefcase, its evaluation of agentic knowledge work, Artificial Analysis says MiniCPM5-2B ranks second in the comparison set at 438, behind Ling 3.0 Tiny (485) and ahead of Granite 4.2 8B (324).

    Combined views

    10.3K

    3 Sources, first seen 23d ago

    Combined views

    10.3K

    3 Sources, first seen 23d ago

    119 likes
    23d ago
    first seen 23d ago
    119 likes
    10 comments
    25 saves
    2 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    10 comments
    25 saves
    2 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    3 Sources

    @ArtificialAnlysAgentic capability is MiniCPM5-2B's edge at this size. On AA-Briefcase, our agentic knowledge work evaluation, its Elo of 438 is second in the set behind Ling 3.0 Tiny (485) and ahead of Granite 4.2 8B (324). The Agentic Index is the weighted average of the three agentic evaluations in the Intelligence Index: AA-Briefcase, GDPval-AA v2 and τ³-Banking.
    @teortaxesTex> level with Qwen3.5 9B (Reasoning, 15, estimated) at roughly 4x its size This is a very good result from @OpenBMB. Qwen 3.5 9B was a strong baseline. Kind of absurd we're still getting such density gains.

    3 Sources

    @ArtificialAnlysAgentic capability is MiniCPM5-2B's edge at this size. On AA-Briefcase, our agentic knowledge work evaluation, its Elo of 438 is second in the set behind Ling 3.0 Tiny (485) and ahead of Granite 4.2 8B (324). The Agentic Index is the weighted average of the three agentic evaluations in the Intelligence Index: AA-Briefcase, GDPval-AA v2 and τ³-Banking.
    @teortaxesTex> level with Qwen3.5 9B (Reasoning, 15, estimated) at roughly 4x its size This is a very good result from @OpenBMB. Qwen 3.5 9B was a strong baseline. Kind of absurd we're still getting such density gains.