Generative reward model aims to scale expert legal review
A developer says their team’s review agents use original outputs, web search and other tools to check its systems’ work, with judgments that correlate strongly with expert lawyer review.
TLDR
A developer says their team built a generative reward model by training agents to approximate partner review of complex legal work. The team combines rubric-based LLM scoring with sampled human preference, but says rubrics miss errors outside predefined criteria, while human reviewers get fatigued and make mistakes at scale. The developer reports that the agents’ judgments correlate strongly with expert lawyer review and expects the approach to help scale human review across model training and production products. They say it will be central to training Tenet 1.5.
Combined views
16.4K
1 Source, first seen 9h ago
Generative reward model aims to scale expert legal review
A developer says their team’s review agents use original outputs, web search and other tools to check its systems’ work, with judgments that correlate strongly with expert lawyer review.