• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    DeepSeek V4 Pro Scores 83.3 on Cybergym Benchmark

    Commentators share benchmark tables for the model across agent tasks with comparisons to Kimi-K3 and other systems.

    T(
    LA
    ZE
    26 Sources, 49d ago, first seen 49d ago

    TLDR

    Posts on X report DeepSeek-V4-Pro-0813 reaching 83.3 on Cybergym and 87.9 on Terminal Bench 2.1. Tables place it ahead of some models including Opus-4.8 in several agent metrics while trailing Kimi-K3 in roughly half the evaluations shown. TeortaxesTex suggests the result may reflect V4-Flash-0731 with pass@2 sampling after internal distillation. Other replies note its position above Grok on ProofBench yet below models such as Muse Spark and question further scaling needs.

    Combined views

    554.6K

    26 Sources, first seen 49d ago

    Combined views

    554.6K

    26 Sources, first seen 49d ago

    3.5K likes
    3.5K likes
    204 comments
    310 saves
    77 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Featured Source
    204 comments
    310 saves
    77 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    26 Sources

    @zephyr_z983.3 on CyberGem Dayum
    @scaling01DeepSeek-V4-Pro Benchmarks still behind Kimi-K3 in ~half of them
    @teortaxesTexWell that's a nice thing to wake up to Scores of "V4 GA" are exactly as expected. A bit weaker than K3, with even less fucks given about sandbagging in Nasty Domains. But I think the gap to 0731 suggests they need more than scale. They need another round of RL env progress.
    @xlr8harder@teortaxesTex Is this plausibly compute access biting RL at scale?

    26 Sources

    @zephyr_z983.3 on CyberGem Dayum
    @scaling01DeepSeek-V4-Pro Benchmarks still behind Kimi-K3 in ~half of them
    @teortaxesTexWell that's a nice thing to wake up to Scores of "V4 GA" are exactly as expected. A bit weaker than K3, with even less fucks given about sandbagging in Nasty Domains. But I think the gap to 0731 suggests they need more than scale. They need another round of RL env progress.
    @xlr8harder@teortaxesTex Is this plausibly compute access biting RL at scale?