Guardrails removed spam, off-topic, unclear, or duplicate replies.
Ask a question below.
Published answers will appear here.
illustrative example: the software subset of the ECI comes with Mechanize's GBAEval, where GLM-5.2 scores 0, likely due to it being text-only (so it cannot verify its outputs). excluding this bench gives GLM-5.2 a huge jump. some of the top models don't have GBAEval scores
not a critique of the ECI per se, thats just how IRT works. but it is important to know these things to accurately interpret the results. the error bars are also massive for all models, so saying that one model is strictly superior cause it got one or two more pts is silly
Guardrails removed spam, off-topic, unclear, or duplicate replies.
Ask a question below.
Published answers will appear here.