On Jev's most confident errors, 96.0% of LLM verdicts reportedly repeat its wrong answer
A user describing their experiments says tuned cascades of AI judges offered limited gains over the best single judge. Combining judges into ensembles did not help much either, they say, because the errors were correlated.
TLDR
A user reports that, on Jev's most confident errors, 96.0% of LLM verdicts repeated its wrong answer—versus about 50% if the errors were independent. They say a judge cascade gained at most 1.5 points over the best single judge with cross-fitted confidence thresholds, and at most 2.0 with oracle thresholds. Judge ensembles also offered little help because their errors were correlated, according to the account. The user connects this result to work by Kim et al. examining more than 350 models, which they say found highly correlated errors among larger, more accurate models—even across providers.
Combined views
1.3K
2 Sources, first seen 7h ago
On Jev's most confident errors, 96.0% of LLM verdicts reportedly repeat its wrong answer
A user describing their experiments says tuned cascades of AI judges offered limited gains over the best single judge. Combining judges into ensembles did not help much either, they say, because the errors were correlated.
TLDR
A user reports that, on Jev's most confident errors, 96.0% of LLM verdicts repeated its wrong answer—versus about 50% if the errors were independent. They say a judge cascade gained at most 1.5 points over the best single judge with cross-fitted confidence thresholds, and at most 2.0 with oracle thresholds. Judge ensembles also offered little help because their errors were correlated, according to the account. The user connects this result to work by Kim et al. examining more than 350 models, which they say found highly correlated errors among larger, more accurate models—even across providers.