Many users mocked inconsistent AI benchmark evaluations, dismissing them as subjective or flawed whenever preferred models underperform.
Based on 5 visible X reactions from 16 accounts; directional sample.
Ask a question below.
Published answers will appear here.
@yacineMTB ngl the funniest part is watching everyone become an expert in evaluation methodology the second the rankings change
@yacineMTB Scientific method: If my model wins → valid benchmark. If my model loses → flawed methodology😂
@yacineMTB The users of a model I don’t use: foolish lemmings falling for benchmaxxing
@yacineMTB This is how you should read academic papers in general.
the benchmark when it agrees with my priors: good benchmark the benchmark when it disagrees with my priors: bad benchmark
can we skip this and just write a bot that scrapes the sentiment of users on twitter with good priors?
Many users mocked inconsistent AI benchmark evaluations, dismissing them as subjective or flawed whenever preferred models underperform.
Based on 5 visible X reactions from 16 accounts; directional sample.
Ask a question below.
Published answers will appear here.