Salesforce AI Research Shares MLLM Judge Bias Paper
Salesforce AI Research posted an arXiv paper on calibration failures in MLLM judges for culturally ambiguous content.
The official Salesforce AI Research account posted a link to the arXiv paper Jury Duty: Calibration and Orientation Failures in MLLM-as-a-Judge Under Cultural Ambiguity. The work studies how multimodal LLMs fail to calibrate judgments when evaluating culturally ambiguous content. It notes that agreement with human annotations is undefined as a validation metric when the human pool is culturally heterogeneous and introduces the VOIR DIRE framework to examine these issues. Listed authors include Daniel J.S. Lee, Harsh Sharma, and others from the team.
Combined views
604
3 posts, first seen 20h ago
Salesforce AI Research Shares MLLM Judge Bias Paper
Salesforce AI Research posted an arXiv paper on calibration failures in MLLM judges for culturally ambiguous content.
The official Salesforce AI Research account posted a link to the arXiv paper Jury Duty: Calibration and Orientation Failures in MLLM-as-a-Judge Under Cultural Ambiguity. The work studies how multimodal LLMs fail to calibrate judgments when evaluating culturally ambiguous content. It notes that agreement with human annotations is undefined as a validation metric when the human pool is culturally heterogeneous and introduces the VOIR DIRE framework to examine these issues. Listed authors include Daniel J.S. Lee, Harsh Sharma, and others from the team.