• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    ProgramBench Metrics Diverge for AI Models

    Tweet notes Raw Pass Rate on ProgramBench rewards some models and punishes others.

    T(
    1 Source, 26d ago, first seen 26d ago

    TLDR

    Teortaxes posted about a divergence between Almost-Resolved and Raw Pass Rate scores on ProgramBench from ValsAI. The tweet states that Raw Pass Rate rewards DeepSeek and GPT-5.4 to 6 while relatively punishing Kimi, GLM, Qwen, and Fable. It asks how to understand the split and suggests possible factors such as RL versus code knowledge coverage or agentic time horizon. The post includes a photo attachment showing the benchmark data.

    Combined views

    4.1K

    1 Source, first seen 26d ago

    Combined views

    4.1K

    1 Source, first seen 26d ago

    27 likes
    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    27 likes
    2 comments
    7 saves

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    2 comments
    7 saves

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    1 Source

    @teortaxesTexInteresting divergence between Almost-Resolved and Raw Pass Rate on ProgramBench from @ValsAI. How do you understand it? RPR rewards: DeepSeek, GPT-5.4 to 6, and relatively punishes: Kimi, GLM, Qwen, Fable. RL vs code knowledge coverage? Agentic time horizon?

    1 Source

    @teortaxesTexInteresting divergence between Almost-Resolved and Raw Pass Rate on ProgramBench from @ValsAI. How do you understand it? RPR rewards: DeepSeek, GPT-5.4 to 6, and relatively punishes: Kimi, GLM, Qwen, Fable. RL vs code knowledge coverage? Agentic time horizon?