• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
AI
Reaction

Evaluating AI data providers beyond benchmark gains

A post argues labs favor data that fixes production failures over gains confined to a vendor’s benchmarks.

1 Source, 34m ago, first seen 34m ago

TLDR

A post argues that labs test purchased data against internal evaluations and may not renew contracts when gains appear only on a vendor’s benchmark. It says labs also pay for hard-to-replicate expertise, reinforcement-learning environments and graders, private evaluations, speed and exclusivity. To judge a provider, the author suggests checking whether gains transfer to related tasks, whether labs keep buying, and whether customers bring vendors problems to solve.

Combined views

—

1 Source, first seen 34m ago

— likes— comments— saves— reposts

Combined views

—

1 Source, first seen 34m ago

— likes— comments— saves— reposts

Sentiment

Positive——Negative

Summary

Not enough discussion yet.

No sentiment analysis available yet.

Featured Source

Sentiment

Positive——Negative

Summary

Not enough discussion yet.

No sentiment analysis available yet.

1 Source

Gokul Rajaram@gokulrEvaluating AI data providers: Beyond benchmark-maxxing Benchmark-maxxing gets an AI data vendor the first sale. Labs are sophisticated buyers. Every serious lab keeps held-out internal evals the vendor never sees. It runs an ablation on each data purchase and compares the gain to the price. Data that only lifts the vendor’s own benchmark looks like noise in that ablation. The vendor gets a first purchase order and never a second. Labs pay for 5 things, roughly in order of how long the spend lasts. 1. Fixes for failures they see in production. A lab knows where its model breaks for paying users from usage logs and enterprise escalations. A coding agent loses track after 40 steps. A finance workflow makes up a number in the third tab of a model. A vendor who can take a cluster of failures and return data that fixes it will get the next contract too. 2. Expertise the lab can’t generate cheaply. Synthetic data and model-generated traces cover a lot now. They can’t reproduce how a tax attorney or a staff engineer reviewing a 2,000-line diff makes judgment calls. Labs pay for access to these people, and more and more they want the expert’s reasoning and rubric along with the answer. 3. Environments and graders for RL. As post-training moves toward reinforcement learning, labs buy environments: a sandboxed task, a way to check whether it succeeded and a reward signal that’s hard to game. A good environment keeps producing training data long after the vendor delivers it. My read is that this is where the biggest checks are going now. 4. Measurement. Labs buy private, uncontaminated eval sets because public benchmarks have leaked into pretraining data. So the vendors gaming public benchmarks are part of why labs need private ones. 5. Speed and exclusivity. An in-house expert network takes a year to build, and buying one takes a month. A lab will sometimes pay extra for exclusivity just to keep a competitor from training on the same data. If you’re evaluating a data company, 4 questions separate real lift from benchmark-maxxing: 1. Does the gain show up on the lab’s internal evals, or only on evals the vendor built? 2. Does it carry over to related tasks the data didn’t target? 3. What’s net revenue retention by lab? If each lab buys once, that’s benchmark-maxxing. Growing contracts mean the data works. 4. Who defines the problem? If labs bring their failures to the vendor, that’s pull. If the vendor shows up with a benchmark and a pitch, that’s push. The AI data companies that last will be the ones labs bring their problems to and that fix production model failures or help them hill climb internal evals, not just external benchmarks.34m
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI

    1 Source

    Gokul Rajaram@gokulrEvaluating AI data providers: Beyond benchmark-maxxing Benchmark-maxxing gets an AI data vendor the first sale. Labs are sophisticated buyers. Every serious lab keeps held-out internal evals the vendor never sees. It runs an ablation on each data purchase and compares the gain to the price. Data that only lifts the vendor’s own benchmark looks like noise in that ablation. The vendor gets a first purchase order and never a second. Labs pay for 5 things, roughly in order of how long the spend lasts. 1. Fixes for failures they see in production. A lab knows where its model breaks for paying users from usage logs and enterprise escalations. A coding agent loses track after 40 steps. A finance workflow makes up a number in the third tab of a model. A vendor who can take a cluster of failures and return data that fixes it will get the next contract too. 2. Expertise the lab can’t generate cheaply. Synthetic data and model-generated traces cover a lot now. They can’t reproduce how a tax attorney or a staff engineer reviewing a 2,000-line diff makes judgment calls. Labs pay for access to these people, and more and more they want the expert’s reasoning and rubric along with the answer. 3. Environments and graders for RL. As post-training moves toward reinforcement learning, labs buy environments: a sandboxed task, a way to check whether it succeeded and a reward signal that’s hard to game. A good environment keeps producing training data long after the vendor delivers it. My read is that this is where the biggest checks are going now. 4. Measurement. Labs buy private, uncontaminated eval sets because public benchmarks have leaked into pretraining data. So the vendors gaming public benchmarks are part of why labs need private ones. 5. Speed and exclusivity. An in-house expert network takes a year to build, and buying one takes a month. A lab will sometimes pay extra for exclusivity just to keep a competitor from training on the same data. If you’re evaluating a data company, 4 questions separate real lift from benchmark-maxxing: 1. Does the gain show up on the lab’s internal evals, or only on evals the vendor built? 2. Does it carry over to related tasks the data didn’t target? 3. What’s net revenue retention by lab? If each lab buys once, that’s benchmark-maxxing. Growing contracts mean the data works. 4. Who defines the problem? If labs bring their failures to the vendor, that’s pull. If the vendor shows up with a benchmark and a pitch, that’s push. The AI data companies that last will be the ones labs bring their problems to and that fix production model failures or help them hill climb internal evals, not just external benchmarks.34m
    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet