An AI agent can produce a convincing final reply even when the work behind it failed. An interactive Jev-as-a-Judge guide demonstrates a different approach: give TypeSafe AI's decision model the full record of the run, not just the answer.
The guide calls that record a trajectory. It includes the original request, the rules the agent should follow, every tool call, tool results such as errors or timeouts, and the final reply. Jev chooses among predefined answers and attaches a probability to each verdict.
The refund demo shows why the trace matters
The guide's playground compares similar final replies across a valid refund, an expired order and an unknown outcome. In its example policy, an 80% threshold determines whether a run passes, fails or lands in a middle band for review.
That routing still needs testing. The guide reports that a compliant escalation response for an unknown outcome scored about 50% in its tests and went to review. The judge also produced a false positive in the guide's tests: a run that skipped the order lookup entirely still scored about 86% likely correct even though the agent never checked eligibility.
Some checks belong in code
The guide recommends collecting realistic trajectories, labeling a sample with people, comparing Jev's verdicts with those labels and rerunning the same examples to test stability. It also advises escalating uncertain or high-risk cases to a stronger judge or a person.
For requirements with one exact answer, the guide argues that code may be a better fit. A rule requiring the agent to call lookup_order before issue_refund, for example, could have caught the missed lookup directly.