Anirudh Goyal Introduces AI Research Preference Models
The models rank candidate experiments using prior results to decide which ones to run next.
Anirudh Goyal, Senior Staff Research Scientist at Google DeepMind, posted about AI Research Preference Models. The work targets the gap between rapid idea generation by research agents and the high cost of running experiments on GPUs. RPMs reframe selection as a ranking task over candidate solutions and research history rather than direct score prediction. An inference-only version uses a frozen LLM to rank plans and code. Tests across 20 ML research tasks with a Qwen3.6-27B backbone produced higher average normalized scores than a no-RPM baseline. An agentic RPM variant performed best among the tested approaches.
Combined views
10.5K
8 posts, first seen 1d ago