Jev’s potential role in AI evaluations and agent decisions
A user who gained access to Jev wants to test it for yes-or-no evaluations, using probabilities with chosen cutoffs instead of relying on an LLM as a judge.
TLDR
A user who gained access to Jev predicts it could replace many uses of LLMs as judges for binary decisions. Their proposed approach: give Jev context and questions, then apply cutoffs to the probabilities it returns. They also see potential for choosing models or tools, deciding whether to continue or retry, checking output quality and flagging work for human review. The idea is to leave generation to the LLM while Jev handles surrounding judgment calls; the user is describing intended uses, not test results.
Combined views
104.1K
13 Sources, first seen 21h ago
Jev’s potential role in AI evaluations and agent decisions
A user who gained access to Jev wants to test it for yes-or-no evaluations, using probabilities with chosen cutoffs instead of relying on an LLM as a judge.