CritPt Benchmark Scores Questioned by AI Commentator
@scaling01 highlights flaws where scores lack sense and fail to improve with reasoning.
TLDR
A tweet by @scaling01 states that CritPt has flaws. The author notes the scores did not make much sense and were not increasing much with extra reasoning. The post comes from a pseudonymous AI commentator who runs LisanBench, a custom LLM reasoning benchmark, and frequently posts technical analysis of model scaling and capabilities. It includes a photo attachment and appears among public posts on X.
Combined views
8.1K
1 Source, first seen 29d ago