LAB benchmark documents allegedly reveal “planted issues” to models
A user points to speaker notes stating “THIS IS THE MOST CRITICAL APPENDIX FOR PLANTED ISSUES,” arguing that these disclosures can skew LAB’s issue-spotting scores.
TLDR
A user analyzing model behavior says some LAB case documents directly disclose the “planted issues” targeted by scoring rubrics. They cite speaker notes in the draft-diligence-summary-memo task that flag an appendix before listing those issues. According to the account, some models fixated on the disclosed issues, often at the expense of broader issue spotting. Others largely ignored the cues, analyzed the documents independently, and frequently scored lower despite producing analyses the user judged more accurate to the documents. The user argues that these disclosures can function like an answer key, undermining LAB’s validity as a measure of issue spotting.
Combined views
7
1 Source, first seen 2d ago
LAB benchmark documents allegedly reveal “planted issues” to models
A user points to speaker notes stating “THIS IS THE MOST CRITICAL APPENDIX FOR PLANTED ISSUES,” arguing that these disclosures can skew LAB’s issue-spotting scores.
TLDR
A user analyzing model behavior says some LAB case documents directly disclose the “planted issues” targeted by scoring rubrics. They cite speaker notes in the draft-diligence-summary-memo task that flag an appendix before listing those issues. According to the account, some models fixated on the disclosed issues, often at the expense of broader issue spotting. Others largely ignored the cues, analyzed the documents independently, and frequently scored lower despite producing analyses the user judged more accurate to the documents. The user argues that these disclosures can function like an answer key, undermining LAB’s validity as a measure of issue spotting.