JEV reportedly matches commercial LLMs on binary decisions, but their errors overlap
EmergentMind describes a paper testing JEV, a rubric judge that returns probabilities over fixed answers instead of generating text. The post says that when JEV makes a confident error, commercial LLM judges make the exact same mistake 96% of the time.
TLDR
EmergentMind says JEV evaluates all criteria for a submission in one request and matches commercial LLMs on binary decisions. But when JEV is confidently wrong, the post says, those LLM judges make the exact same mistake 96% of the time. It also says automated models assign lower grades than human raters and argues that shared errors limit the value of fallback judge cascades.
Combined views
876
2 Sources, first seen 10h ago
