• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI
    Reaction

    Jev-as-a-Judge checks the work behind an AI agent's final answer

    The interactive guide feeds Jev full trajectories, then routes uncertain verdicts to review instead of trusting polished replies.

    elvisEL
    DAIR.AIDA
    3 Sources, 2h ago, first seen 2h ago

    TLDR

    Jev-as-a-Judge is a tutorial for evaluating an AI agent’s full trajectory, including its rules, tool calls, results and final reply. In the guide’s tests, Jev still rated a run that skipped an order lookup about 86% likely correct. The guide recommends calibrating against human labels, checking repeat-run stability and routing uncertain or high-risk cases to stronger judges or people.

    Combined views

    17.1K

    3 Sources, first seen 2h ago

    Combined views

    17.1K

    3 Sources, first seen 2h ago

    130 likes

    Useful links

    arXiv.org

    JEV-as-a-Judge: Accept When Confident, Escalate When Unsure

    AI Paper Slop · YouTube

    JEV-as-a-Judge: Accept When Confident, Escalate When Unsure (Sep 2026)

    Made with Jev

    jev-as-judge — built with Jev

    Prompt Engineer 48 · YouTube

    LangChain + Jev Full Tutorial: Routing, Guardrails & Evals for $0.00002 a Decision

    OpenRouter Blog

    Jev vs LLM-as-a-Judge

    Sam Witteveen · YouTube

    Using Jev In Your Agent Harness

    LangChain Blog

    Can Jev Be a Better Agent Evaluator?
    130 likes
    25 comments
    164 saves
    17 reposts
    25 comments
    164 saves
    17 reposts

    An AI agent can produce a convincing final reply even when the work behind it failed. An interactive Jev-as-a-Judge guide demonstrates a different approach: give TypeSafe AI's decision model the full record of the run, not just the answer.

    Featured Source

    The guide calls that record a trajectory. It includes the original request, the rules the agent should follow, every tool call, tool results such as errors or timeouts, and the final reply. Jev chooses among predefined answers and attaches a probability to each verdict.

    The refund demo shows why the trace matters

    The guide's playground compares similar final replies across a valid refund, an expired order and an unknown outcome. In its example policy, an 80% threshold determines whether a run passes, fails or lands in a middle band for review.

    That routing still needs testing. The guide reports that a compliant escalation response for an unknown outcome scored about 50% in its tests and went to review. The judge also produced a false positive in the guide's tests: a run that skipped the order lookup entirely still scored about 86% likely correct even though the agent never checked eligibility.

    Some checks belong in code

    The guide recommends collecting realistic trajectories, labeling a sample with people, comparing Jev's verdicts with those labels and rerunning the same examples to test stability. It also advises escalating uncertain or high-risk cases to a stronger judge or a person.

    For requirements with one exact answer, the guide argues that code may be a better fit. A rule requiring the agent to call lookup_order before issue_refund, for example, could have caught the missed lookup directly.

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Useful Links

    arXiv.org

    JEV-as-a-Judge: Accept When Confident, Escalate When Unsure

    Made with Jev

    jev-as-judge — built with Jev

    OpenRouter Blog

    Jev vs LLM-as-a-Judge

    LangChain Blog

    Can Jev Be a Better Agent Evaluator?

    Related Videos

    • JEV-as-a-Judge: Accept When Confident, Escalate When Unsure (Sep 2026)AI Paper Slop · YouTube
    • LangChain + Jev Full Tutorial: Routing, Guardrails & Evals for $0.00002 a DecisionPrompt Engineer 48 · YouTube
    • Using Jev In Your Agent HarnessSam Witteveen · YouTube
    Today's Rank

    #8

    Today's Rank

    #8

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    3 Sources

    elvis@omarsar0https://x.com/i/article/21072584656305070092h
    DAIR.AI@dair_aiImprove the reliability of agent evaluation with Jev-as-a-Judge.2h

    Useful Links

    arXiv.org

    JEV-as-a-Judge: Accept When Confident, Escalate When Unsure

    Related Videos

    Made with Jev

    jev-as-judge — built with Jev

    OpenRouter Blog

    Jev vs LLM-as-a-Judge

    LangChain Blog

    Can Jev Be a Better Agent Evaluator?
    JEV-as-a-Judge: Accept When Confident, Escalate When Unsure (Sep 2026)AI Paper Slop · YouTube
  • LangChain + Jev Full Tutorial: Routing, Guardrails & Evals for $0.00002 a DecisionPrompt Engineer 48 · YouTube
  • Using Jev In Your Agent HarnessSam Witteveen · YouTube
  • 3 Sources

    elvis@omarsar0https://x.com/i/article/21072584656305070092h
    DAIR.AI@dair_aiImprove the reliability of agent evaluation with Jev-as-a-Judge.2h