How are benchmarks used in model development, and where do they fall short?
SnorkelAI highlights a Frontier Data Summit discussion featuring participants from Laude Institute, Open Athena and Stanford CRFM.
TLDR
SnorkelAI is promoting a Frontier Data Summit discussion about benchmarks’ role in model development and their limitations. It pairs a participant from Laude Institute with an Open Athena technical staff member who also leads research engineering at Stanford CRFM.
How are benchmarks used in model development, and where do they fall short?
SnorkelAI highlights a Frontier Data Summit discussion featuring participants from Laude Institute, Open Athena and Stanford CRFM.
TLDR
SnorkelAI is promoting a Frontier Data Summit discussion about benchmarks’ role in model development and their limitations. It pairs a participant from Laude Institute with an Open Athena technical staff member who also leads research engineering at Stanford CRFM.