• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI
    Reaction

    AI data vendors' staying power through the lens of early ad networks

    A post argues that fast-growing AI data suppliers, like early ad networks, may struggle to become harder to replace.

    GR
    2 Sources, ,

    TLDR

    A post compares AI data vendors with early ad networks, arguing that rapid revenue growth alone may not build durable value if suppliers share experts and labs own the output. It proposes three ways to become harder to replace: own differentiated data supply, become part of labs' workflows, and build training environments that improve using model feedback. The author predicts that vertical specialists will replace large horizontal providers.

    Combined views

    3.5K

    2 Sources, first seen 1h ago

    Combined views

    3.5K

    2 Sources, first seen 1h ago

    25 likes
    1h ago
    first seen 1h ago
    25 likes
    6 comments
    45 saves
    4 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Featured Source
    6 comments
    45 saves
    4 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    2 Sources

    @gokulrDefensibility in Data: Lessons from Ads (Warning: LONG POST) Companies selling data / expertise / RL environments to AI labs are growing at extraordinary speed, with several reaching tens to hundreds of millions in annualized revenue within a year. But venture investors are unsure how to value these companies, given customer concentration, limited recurring revenue, and the constant treadmill of new data needs from the labs, which leads to every "product" (i.e. data feed / RL environment) having a very finite lifetime. I see many parallels with advertising in the early 2000s. Back then, dozens of ad networks sprouted up as the number of websites ballooned (thanks to blogging platforms, cheap hosting and CMSs like WordPress). These networks acted as middlemen; they bought unsold inventory from thousands of small and mid-sized publishers, bundled it into audience segments based on behavioral, vertical or contextual targets, and resold it to advertisers who couldn't negotiate deals site by site. These ad networks solved a real problem, since most advertisers couldn’t manage relationships with thousands of websites, while publishers couldn’t build sales teams to reach every advertiser. Networks made fragmented supply easy to buy. Revenue grew exceedingly fast, and several got to tens to even hundreds of millions in revenue, which back then was very significant. Over time, a problem emerged: many networks had access to the same inventory. Publishers worked with multiple networks to fill their inventory. Advertisers spread budgets across them but struggled to track where their money went. Revenue grew pretty without necessarily making the businesses harder to replace. To solve this fragmentation, ad exchanges such as Right Media, DoubleClick Ad Exchange, and Microsoft's AdECN emerged around 2005–2007. They created marketplaces where buyers and sellers could trade impressions directly, laying the groundwork for real-time bidding and the programmatic advertising ecosystem that dominates today. As these publisher-side ad exchanges (and on the advertiser side, Demand Side Platforms) matured, access to inventory became less differentiated. The question became: what does this network contribute that the buyer can’t get elsewhere? Ultimately, most ad networks shut down, got swallowed up and there are very few, I can think of that generated much equity value. Most of the equity value accumulated to the exchanges and the DSPs that centralized buying and selling. AI data companies face a similar question. Today, AI labs need experts, demonstrations, evaluations and increasingly complex tasks. Finding contributors, verifying their capabilities and delivering quality at speed is valuable work. But if the same experts work for five vendors, the lab specifies the task, and the lab owns the output, what is the supplier accumulating? The lab gets a better model. The supplier gets paid and starts looking for the next project. That can be a very good revenue business. Unfortunately, this doesn't mean it's building durable equity value or is a true technology platform. I see three ways to build durable equity in this space: 1. Own differentiated supply: Google and Facebook were the biggest winners in advertising because they both owned their inventory and didn't rent it. Data vendors should similarly have a product, community or exclusive relationship that continually generates useful data. The ability to produce the next valuable dataset matters more than possession of the last one. 2. Become part of the customer’s workflow: Doubleclick (Google paid 1% of its market cap for it) was deeply embedded in both agency and publisher workflows. The Trade Desk ($10B public company) is the platform used by many agencies to buy media. Similarly, the data company needs to become indispensable to the lab: move one step earlier in the funnel and help the lab identify model weaknesses, design tasks, evaluate results and decide what data to buy next. The deeper you participate in those decisions, the harder you are to replace. 3. Build self-improving environments: Invest in technology that turns model failures into new training tasks, verifies outcomes and adjusts difficulty as models improve. Each training cycle should make the environment better at identifying and teaching what the model still gets wrong. The durable asset is the software that generates and evaluates the next useful task. That requires access to performance feedback; if all the learning stays inside the lab, the supplier keeps producing data without building a compounding technology advantage. This is something you can absolutely charge a recurring fee for. #1-3 collectively also likely mean that data vendors should consider specializing in a handful of verticals where they have unique advantages. I believe that large horizontal data provider will be replaced by vertical specialists who own differentiated supply for that vertical, run customer workflows for that vertical and build self-improving RL environments. AppLovin ($150B market cap) is the largest stand alone ad-tech company today. They do exactly ONE THING - drive app installs - and they do it damn well. They own the end to end value chain and have built incredibly differentiated supply as well as a self-improving loop that delivers downloads at ever improving cost-per-install. If they had tried to do anything more, they would likely not have seen similar success. The question every data founder should ask themselves: after delivering your first $100M of revenue, what will you own that makes the next $100M easier to win and harder for a competitor to take? The answer should be more than a larger roster of experts or stronger relationships with the labs. It should be unique supply, deeply embedded workflows, or technology that gets better with every training cycle. The demand boom gives you the revenue to build those assets. Whether you do will determine how much equity value remains as sourcing data becomes easier and (inevitably) commoditized. PS: If you're a founder working on the above, please DM or email me at gokulr at gmail. I would LOVE to speak with you and brainstorm!

    2 Sources

    @gokulrDefensibility in Data: Lessons from Ads (Warning: LONG POST) Companies selling data / expertise / RL environments to AI labs are growing at extraordinary speed, with several reaching tens to hundreds of millions in annualized revenue within a year. But venture investors are unsure how to value these companies, given customer concentration, limited recurring revenue, and the constant treadmill of new data needs from the labs, which leads to every "product" (i.e. data feed / RL environment) having a very finite lifetime. I see many parallels with advertising in the early 2000s. Back then, dozens of ad networks sprouted up as the number of websites ballooned (thanks to blogging platforms, cheap hosting and CMSs like WordPress). These networks acted as middlemen; they bought unsold inventory from thousands of small and mid-sized publishers, bundled it into audience segments based on behavioral, vertical or contextual targets, and resold it to advertisers who couldn't negotiate deals site by site. These ad networks solved a real problem, since most advertisers couldn’t manage relationships with thousands of websites, while publishers couldn’t build sales teams to reach every advertiser. Networks made fragmented supply easy to buy. Revenue grew exceedingly fast, and several got to tens to even hundreds of millions in revenue, which back then was very significant. Over time, a problem emerged: many networks had access to the same inventory. Publishers worked with multiple networks to fill their inventory. Advertisers spread budgets across them but struggled to track where their money went. Revenue grew pretty without necessarily making the businesses harder to replace. To solve this fragmentation, ad exchanges such as Right Media, DoubleClick Ad Exchange, and Microsoft's AdECN emerged around 2005–2007. They created marketplaces where buyers and sellers could trade impressions directly, laying the groundwork for real-time bidding and the programmatic advertising ecosystem that dominates today. As these publisher-side ad exchanges (and on the advertiser side, Demand Side Platforms) matured, access to inventory became less differentiated. The question became: what does this network contribute that the buyer can’t get elsewhere? Ultimately, most ad networks shut down, got swallowed up and there are very few, I can think of that generated much equity value. Most of the equity value accumulated to the exchanges and the DSPs that centralized buying and selling. AI data companies face a similar question. Today, AI labs need experts, demonstrations, evaluations and increasingly complex tasks. Finding contributors, verifying their capabilities and delivering quality at speed is valuable work. But if the same experts work for five vendors, the lab specifies the task, and the lab owns the output, what is the supplier accumulating? The lab gets a better model. The supplier gets paid and starts looking for the next project. That can be a very good revenue business. Unfortunately, this doesn't mean it's building durable equity value or is a true technology platform. I see three ways to build durable equity in this space: 1. Own differentiated supply: Google and Facebook were the biggest winners in advertising because they both owned their inventory and didn't rent it. Data vendors should similarly have a product, community or exclusive relationship that continually generates useful data. The ability to produce the next valuable dataset matters more than possession of the last one. 2. Become part of the customer’s workflow: Doubleclick (Google paid 1% of its market cap for it) was deeply embedded in both agency and publisher workflows. The Trade Desk ($10B public company) is the platform used by many agencies to buy media. Similarly, the data company needs to become indispensable to the lab: move one step earlier in the funnel and help the lab identify model weaknesses, design tasks, evaluate results and decide what data to buy next. The deeper you participate in those decisions, the harder you are to replace. 3. Build self-improving environments: Invest in technology that turns model failures into new training tasks, verifies outcomes and adjusts difficulty as models improve. Each training cycle should make the environment better at identifying and teaching what the model still gets wrong. The durable asset is the software that generates and evaluates the next useful task. That requires access to performance feedback; if all the learning stays inside the lab, the supplier keeps producing data without building a compounding technology advantage. This is something you can absolutely charge a recurring fee for. #1-3 collectively also likely mean that data vendors should consider specializing in a handful of verticals where they have unique advantages. I believe that large horizontal data provider will be replaced by vertical specialists who own differentiated supply for that vertical, run customer workflows for that vertical and build self-improving RL environments. AppLovin ($150B market cap) is the largest stand alone ad-tech company today. They do exactly ONE THING - drive app installs - and they do it damn well. They own the end to end value chain and have built incredibly differentiated supply as well as a self-improving loop that delivers downloads at ever improving cost-per-install. If they had tried to do anything more, they would likely not have seen similar success. The question every data founder should ask themselves: after delivering your first $100M of revenue, what will you own that makes the next $100M easier to win and harder for a competitor to take? The answer should be more than a larger roster of experts or stronger relationships with the labs. It should be unique supply, deeply embedded workflows, or technology that gets better with every training cycle. The demand boom gives you the revenue to build those assets. Whether you do will determine how much equity value remains as sourcing data becomes easier and (inevitably) commoditized. PS: If you're a founder working on the above, please DM or email me at gokulr at gmail. I would LOVE to speak with you and brainstorm!