Upwork Researchers Release UPHELD Conversation Benchmark
DAIR.AI highlights a new benchmark using scripted human dialogues to test LLM judges.
TLDR
DAIR.AI posted about a paper from Upwork researchers titled Evaluating Language Models in Realistic Conversational Contexts. The work introduces UPHELD, a benchmark built from hundreds of complete human-to-human dialogues written by professional script writers. The account notes that LLM judges can disagree with experts and presents the benchmark as one way to examine those differences. The paper appears on arXiv with authors Ilija Subasic, Andrew Rabinovich, and Zhao Chen.
Combined views
9K
1 Source, first seen 32d ago