• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
AI

Jev reportedly beats LLM judges on speed and cost in a narrow test

LangChain says TypeSafe AI’s Jev returns structured answers rather than generating text. In its test, Jev’s quality-score variance was 92–913 times lower than that of GPT-5.6 Luna, Terra and Claude Sonnet 4.6.

Harrison ChaseHC
elvisEL
LangChainLA
9 Sources, 20d ago, first seen 20d ago

TLDR

LangChain tested TypeSafe AI’s Jev against LLM judges on accuracy, repeatability, latency and cost. It reports that Jev was the fastest and cheapest evaluator in the comparison, averaging 0.44 seconds and $0.00035 per call. Jev also produced more consistent continuous scores, with substantially lower quality-score variance. LangChain describes the results as promising but early, emphasizing the test’s narrow scope.

Combined views

126.1K

9 Sources, first seen 20d ago

911 likes127 comments681 saves111 reposts

Combined views

126.1K

9 Sources, first seen 20d ago

911 likes127 comments681 saves111 reposts

Sentiment

Positive——Negative

Summary

Not enough discussion yet.

No sentiment analysis available yet.

Sentiment

Positive——Negative

Summary

Not enough discussion yet.

No sentiment analysis available yet.

9 Sources

Punit Arani@punit_arani5x faster than LLMs while being 13% cheaper than GPT 5.6 Luna, 88% cheaper than GPT 5.6 terra and 99% cheaper than Sonnet 4.6 is insane Has anyone else used Jev for evals?20d
Harrison Chase@hwchase17RT @punit_arani: 5x faster than LLMs while being 13% cheaper than GPT 5.6 Luna, 88% cheaper than GPT 5.6 terra and 99% cheaper than Sonnet…20d
LangChain@LangChainJev-as-a-judge is now available in LangSmith. ✅ Score every production trace instead of a sample. ✅ Check more criteria per trace without the cost climbing. ✅ Catch safety or security issues fast enough to trigger an automated response. Give it a try and let us know what you think! https://www.langchain.com/blog/jev-is-now-available-in-langsmith-evals19d
Viv@Vtrivedy10if there’s a tool to better understand your data at scale you best believe we’re making it super easy to use it in LangSmith🫡 use Jev for large scale trace mining today, it’s so cheap that you probably don’t even need to sub-sample in LangSmith Gateway you can also try other OSS Jev variants to see what’s best for your tasks - understand every piece of trace data - turn it into evals & environments - build continuously improving agents from those evals19d
elvis@omarsar0Don't sleep on using Jev-as-a-Judge for agent evaluation. This is one of the most impressive Jev use cases I have found so far. Jev is a natural fit as a Judge, but it doesn't mean you use it everywhere. Similarly, you shouldn't use frontier models for evals everywhere. I'm running lots of tests on this atm, but early results point to an optimized flow (balancing accuracy and cost) that combines Jev and frontier models. Concretely, use Jev in high-confidence situations, and escalate to a frontier model (GPT-6 or Opus 5.5) in low-confidence verdicts. Entire write-up coming soon. Let me know if you have questions as I build the full guide.16d
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI

    9 Sources

    Punit Arani@punit_arani5x faster than LLMs while being 13% cheaper than GPT 5.6 Luna, 88% cheaper than GPT 5.6 terra and 99% cheaper than Sonnet 4.6 is insane Has anyone else used Jev for evals?20d
    Harrison Chase@hwchase17RT @punit_arani: 5x faster than LLMs while being 13% cheaper than GPT 5.6 Luna, 88% cheaper than GPT 5.6 terra and 99% cheaper than Sonnet…20d
    LangChain@LangChainJev-as-a-judge is now available in LangSmith. ✅ Score every production trace instead of a sample. ✅ Check more criteria per trace without the cost climbing. ✅ Catch safety or security issues fast enough to trigger an automated response. Give it a try and let us know what you think! https://www.langchain.com/blog/jev-is-now-available-in-langsmith-evals19d
    Viv@Vtrivedy10if there’s a tool to better understand your data at scale you best believe we’re making it super easy to use it in LangSmith🫡 use Jev for large scale trace mining today, it’s so cheap that you probably don’t even need to sub-sample in LangSmith Gateway you can also try other OSS Jev variants to see what’s best for your tasks - understand every piece of trace data - turn it into evals & environments - build continuously improving agents from those evals19d
    elvis@omarsar0Don't sleep on using Jev-as-a-Judge for agent evaluation. This is one of the most impressive Jev use cases I have found so far. Jev is a natural fit as a Judge, but it doesn't mean you use it everywhere. Similarly, you shouldn't use frontier models for evals everywhere. I'm running lots of tests on this atm, but early results point to an optimized flow (balancing accuracy and cost) that combines Jev and frontier models. Concretely, use Jev in high-confidence situations, and escalate to a frontier model (GPT-6 or Opus 5.5) in low-confidence verdicts. Entire write-up coming soon. Let me know if you have questions as I build the full guide.16d
    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet