Reply Challenges Solvability Claims on HLE Benchmark
Teortaxes questions earlier assertions about task difficulty in the benchmark.
TLDR
Teortaxes, a commentator focused on DeepSeek, responded to xeophon about solvability in the HLE benchmark. The reply suggests that assertions of high solvability rates might be overstated. It mentions that in areas such as bio, around half the tasks appear unsolvable. The post further indicates that complete solvability is not essential based on recent insights. This contributes to the visible discussion among replies on the platform regarding what the benchmark actually measures.
Combined views
1.8K
1 Source, first seen 32d ago