JEV judge cascade reportedly keeps 99% of GPT-6's accuracy at about 57% of its fee
DAIR.AI describes a paper testing a two-model approach on 510 held-out preference pairs: accept confident verdicts from TypeSafe AI's JEV judge and send uncertain calls to GPT-6 Astra.
TLDR
According to DAIR.AI's summary of the Jev-as-a-Judge paper, combining JEV's confident verdicts with GPT-6 Astra for uncertain decisions retained 99% of GPT-6's accuracy at about 57% of its fee on 510 held-out preference pairs. JEV returns verdicts and label probabilities without reasoning text. The summary notes a limitation: JEV's accuracy gap versus GPT-6 grew to 9–20 points on tasks requiring derivation checks or rejection of elaborately written wrong answers. It also says the authors recommend setting the escalation threshold on your own data, because it did not transfer to every fallback model.
