A calibration benchmark for admitting language models to human-computation tasks
A post describes the work as an approach to auditing LLM annotators, similar to qualifying crowd workers.
TLDR
A September 29 post says the work considers auditing LLM annotators much like qualifying crowd workers. The CI ’26 program lists it as “Qualification by Calibration: A Readable Benchmark for Admitting Language Models to Human-Computation Tasks.”
Combined views
86
1 Source, first seen 2h ago
