Some users praise the analysis of SimpleQA benchmark flaws and engagement on ECI improvements, while many others dismiss the leaderboard shifts as gimmicky with meaningless noisy numbers.
Based on 9 visible X reactions from 12 accounts; directional sample.
Ask a question below.
Published answers will appear here.
Score adjustments ranged between +0.13 and -0.75 points.
@scaling01 I called the SimpleQA impact for open models such as GPT-OSS end of last year in private to them like it’s a fundamental thing of IRT, dunno what you are on about
@scaling01 he's been showing the flaws in ECI for months and actively engages with how to improve it. A related post: https://www.interconnects.ai/p/latest-open-artifacts-21-open-model
@scaling01 The difference of Kimi K3 and Gemini 3.1 is like half a point. The numbers are meaningless and noisy
@scaling01 Fair that my impact for modern models was overstated, good that you did the analysis.
@scaling01 You're turning into a gimmick account
@anacreonte_ @scaling01 yeah it's very gimmicky by now
they are coping so hard that they are already making up numbers to discredit ECI
Mean change: +0.21 ECI 27 of 30 change by less than 0.5 ECI
Some users praise the analysis of SimpleQA benchmark flaws and engagement on ECI improvements, while many others dismiss the leaderboard shifts as gimmicky with meaningless noisy numbers.
Based on 9 visible X reactions from 12 accounts; directional sample.
Ask a question below.
Published answers will appear here.