DeepSeek V4 Pro Scores 83.3 on Cybergym Benchmark
Commentators share benchmark tables for the model across agent tasks with comparisons to Kimi-K3 and other systems.
TLDR
Posts on X report DeepSeek-V4-Pro-0813 reaching 83.3 on Cybergym and 87.9 on Terminal Bench 2.1. Tables place it ahead of some models including Opus-4.8 in several agent metrics while trailing Kimi-K3 in roughly half the evaluations shown. TeortaxesTex suggests the result may reflect V4-Flash-0731 with pass@2 sampling after internal distillation. Other replies note its position above Grok on ProofBench yet below models such as Muse Spark and question further scaling needs.
Combined views
554.6K
26 Sources, first seen 49d ago
