• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Upwork Researchers Release UPHELD Conversation Benchmark

    DAIR.AI highlights a new benchmark using scripted human dialogues to test LLM judges.

    DA
    1 Source, 32d ago, first seen 32d ago

    TLDR

    DAIR.AI posted about a paper from Upwork researchers titled Evaluating Language Models in Realistic Conversational Contexts. The work introduces UPHELD, a benchmark built from hundreds of complete human-to-human dialogues written by professional script writers. The account notes that LLM judges can disagree with experts and presents the benchmark as one way to examine those differences. The paper appears on arXiv with authors Ilija Subasic, Andrew Rabinovich, and Zhao Chen.

    Combined views

    9K

    1 Source, first seen 32d ago

    Combined views

    9K

    1 Source, first seen 32d ago

    83 likes
    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    83 likes
    16 comments
    61 saves
    8 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    16 comments
    61 saves
    8 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    1 Source

    @dair_ai// Your LLM judge disagrees with the experts // LLM Judges can be tricky to build. Here is an interesting showcasing why: There propose a reference-full benchmark of hundreds of complete human-to-human dialogues written by professional script writers, with realistic turn densities and more than 36,000 per-turn human annotations across over 30,000 expert-generated turns. Conversational evaluation frameworks were mostly built for summarization, translation and short-form QA, and the metrics themselves are often derived and validated on synthetic data rather than human dialogue. Tested against expert judgment at this scale, both classical automatic metrics and reference-free LLM-as-a-judge approaches turn out to be unreliable. Their Mixture-of-Judges framework combines multiple evaluative signals and recovers roughly 30 percent better correlation with human assessment. Paper: https://arxiv.org/abs/2608.26131 Chat with Paper: https://academy.dair.ai/papers/evaluating-language-models-in-realistic-conversational-contexts-2608.26131

    1 Source

    @dair_ai// Your LLM judge disagrees with the experts // LLM Judges can be tricky to build. Here is an interesting showcasing why: There propose a reference-full benchmark of hundreds of complete human-to-human dialogues written by professional script writers, with realistic turn densities and more than 36,000 per-turn human annotations across over 30,000 expert-generated turns. Conversational evaluation frameworks were mostly built for summarization, translation and short-form QA, and the metrics themselves are often derived and validated on synthetic data rather than human dialogue. Tested against expert judgment at this scale, both classical automatic metrics and reference-free LLM-as-a-judge approaches turn out to be unreliable. Their Mixture-of-Judges framework combines multiple evaluative signals and recovers roughly 30 percent better correlation with human assessment. Paper: https://arxiv.org/abs/2608.26131 Chat with Paper: https://academy.dair.ai/papers/evaluating-language-models-in-realistic-conversational-contexts-2608.26131