A proposal for fully open or fully closed AI evaluations
One post argues that private held-out test sets offer little signal about overall quality while allowing the same synthetic-data pipelines as open evaluations.
TLDR
One post proposes only two types of AI evaluations: fully open source so the community can audit everything, or fully closed with a one-sentence description. The author argues that private held-out sets give little signal about overall quality but still allow the same synthetic-data pipelines as open evaluations.
Combined views
5.7K
1 Source, first seen 1d ago
A proposal for fully open or fully closed AI evaluations
One post argues that private held-out test sets offer little signal about overall quality while allowing the same synthetic-data pipelines as open evaluations.
TLDR
One post proposes only two types of AI evaluations: fully open source so the community can audit everything, or fully closed with a one-sentence description. The author argues that private held-out sets give little signal about overall quality but still allow the same synthetic-data pipelines as open evaluations.