NeurIPS Reviews Often Propose Impractical LLM Experiments
Researchers note NeurIPS reviews often suggest LLM tests too costly to run or refute.
TLDR
Gael Varoquaux states many NeurIPS reviews propose empirical tests whose scale exceeds available budgets, preventing researchers from disproving flawed feedback. In the thread with Dan Roy, the pair discuss prior-fitted networks, traditional tabular methods, and frontier LLMs. Varoquaux maintains LLMs lag conventional machine learning on tables. The exchange illustrates how reviewer demands for large-scale runs create barriers for teams without equivalent compute resources.
NeurIPS Reviews Often Propose Impractical LLM Experiments
Researchers note NeurIPS reviews often suggest LLM tests too costly to run or refute.
TLDR
Gael Varoquaux states many NeurIPS reviews propose empirical tests whose scale exceeds available budgets, preventing researchers from disproving flawed feedback. In the thread with Dan Roy, the pair discuss prior-fitted networks, traditional tabular methods, and frontier LLMs. Varoquaux maintains LLMs lag conventional machine learning on tables. The exchange illustrates how reviewer demands for large-scale runs create barriers for teams without equivalent compute resources.