• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Generative reward model aims to scale expert legal review

    A developer says their team’s review agents use original outputs, web search and other tools to check its systems’ work, with judgments that correlate strongly with expert lawyer review.

    Gabe PereyraGP
    1 Source, 20d ago, first seen 20d ago

    TLDR

    A developer says their team built a generative reward model by training agents to approximate partner review of complex legal work. The team combines rubric-based LLM scoring with sampled human preference, but says rubrics miss errors outside predefined criteria, while human reviewers get fatigued and make mistakes at scale. The developer reports that the agents’ judgments correlate strongly with expert lawyer review and expects the approach to help scale human review across model training and production products. They say it will be central to training Tenet 1.5.

    Combined views

    34.3K

    1 Source, first seen 20d ago

    Combined views

    34.3K

    1 Source, first seen 20d ago

    83 likes
    83 likes
    2 comments
    86 saves
    7 reposts
    2 comments
    86 saves
    7 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    1 Source

    Gabe Pereyra@gabepereyraThe gold standard for evaluating complex legal work is partner review. Partners can cost over $3,000 an hour and associates $1,000, making expert review expensive at scale. Evaluating a model on LAB through expert review alone would cost millions of dollars. In practice, we combine rubric-based LLM scoring with sampled human preference. But rubrics miss errors beyond predefined criteria, and human reviewers get fatigued and make mistakes at scale. We built a generative reward model to bridge this gap by training agents to approximate partner review. We give these agents the original outputs, web search, and other tools to check our systems’ work. We find that their judgments correlate strongly with expert lawyer review. This approach will help us scale human review across model training and production products, and will be central to training Tenet 1.5.20d

    1 Source

    Gabe Pereyra@gabepereyraThe gold standard for evaluating complex legal work is partner review. Partners can cost over $3,000 an hour and associates $1,000, making expert review expensive at scale. Evaluating a model on LAB through expert review alone would cost millions of dollars. In practice, we combine rubric-based LLM scoring with sampled human preference. But rubrics miss errors beyond predefined criteria, and human reviewers get fatigued and make mistakes at scale. We built a generative reward model to bridge this gap by training agents to approximate partner review. We give these agents the original outputs, web search, and other tools to check our systems’ work. We find that their judgments correlate strongly with expert lawyer review. This approach will help us scale human review across model training and production products, and will be central to training Tenet 1.5.20d