• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Harry Stebbings Outlines Two LLM Win Strategies

    20VC founder shares his view on paths for large language model providers.

    AA
    HS
    JL
    6 Sources, 29d ago, first seen 29d ago

    TLDR

    Harry Stebbings, founder and host of the 20VC podcast and founder of the 20VC VC fund, posted a quote on the two paths large language models will win. Model providers can dominate the platform era by selling inference or shift upward to become application-layer companies with strong models. He stated that Anthropic seems to be following the application path. The post includes a 25-second video attachment of Eno Reye.

    Combined views

    81.1K

    6 Sources, first seen 29d ago

    Combined views

    81.1K

    6 Sources, first seen 29d ago

    185 likes
    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    185 likes
    67 comments
    198 saves
    17 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    67 comments
    198 saves
    17 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    6 Sources

    @HarryStebbingsThe two ways that large language models will win “If you are a model provider, you are basically looking to dominate the platform era and get really good at selling inference. Or you want to move up to become an application-layer company that has really good models. Anthropic seems to be following the application path, while OpenAI seems to be dipping its toes in both.” @EnoReyes How do you think about this @charlespacker @swyx @bennstancil @jaminball
    @jerryjliu0Vendor-neutral startups are uniquely positioned to solve any given task better than the frontier labs. They have access to the full set of open-weight and frontier models. They can still use the frontier models where needed, but can optimize the harness e2e for any given task. They can even posttrain the model itself because they have the domain expertise to create specialized evals.
    @matanSFRT @HarryStebbings: What is required to get the most out of models? “To take the most advantage out of models, you need to do something a…
    @gokulr100% agree with @EnoReyes Stateful intelligence allocation (vs token routing) is the way to optimize model use.
    @AustenThis is an excellent question. Why aren’t AI models just training themselves already? They theoretically can, and kind of are, but they don’t have the data/evals/gyms required to do so. A short reading list: Bottleneck is the environment, not compute https://medium.com/@shuchaobi/ais-next-bottleneck-isn-t-compute-it-s-the-environment-0eaec31888f5 RL envs cost real money ($20k–$300k/env, and this is for simulated ones which are just kinda crappy IMO) https://epoch.ai/gradient-updates/state-of-rl-envs We’re running out of human text https://epoch.ai/publications/will-we-run-out-of-data-limits-of-llm-scaling-based-on-human-generated-data Eval is the bottleneck https://ysymyth.github.io/The-Second-Half/ Verifier’s law https://www.jasonwei.net/blog/asymmetry-of-verification-and-verifiers-law What labs buy: Foody on RL envs https://www.youtube.com/watch?v=a00xIn5kwhM Economy as RL environment machine https://www.mercor.com/blog/the-economy-will-become-an-rl-environment-machine/ APEX-Agents generalization https://www.mercor.com/blog/generalization-results-from-training-on-the-apex-agents-dev-set/ Etna: ~$1B/yr on external data, supply-constrained Surge Tuesday (can it get through a workday?) https://surgehq.ai/blog/tuesday-frontier-work-index Dario: task/process distribution, not more web text https://www.dwarkesh.com/p/dario-amodei-2 OSWorld 2.0 (~20% on long workflows) https://osworld-v2.xlang.ai/ https://arxiv.org/abs/2606.29537 SWE-bench = ticket + repo + tests https://www.swebench.com Karpathy: sucking supervision through a straw https://www.dwarkesh.com/p/andrej-karpathy Ilya: peak data / one internet https://www.reuters.com/technology/artificial-intelligence/ai-with-reasoning-power-will-be-less-predictable-ilya-sutskever-says-2024-12-14/

    6 Sources

    @HarryStebbingsThe two ways that large language models will win “If you are a model provider, you are basically looking to dominate the platform era and get really good at selling inference. Or you want to move up to become an application-layer company that has really good models. Anthropic seems to be following the application path, while OpenAI seems to be dipping its toes in both.” @EnoReyes How do you think about this @charlespacker @swyx @bennstancil @jaminball
    @jerryjliu0Vendor-neutral startups are uniquely positioned to solve any given task better than the frontier labs. They have access to the full set of open-weight and frontier models. They can still use the frontier models where needed, but can optimize the harness e2e for any given task. They can even posttrain the model itself because they have the domain expertise to create specialized evals.
    @matanSFRT @HarryStebbings: What is required to get the most out of models? “To take the most advantage out of models, you need to do something a…
    @gokulr100% agree with @EnoReyes Stateful intelligence allocation (vs token routing) is the way to optimize model use.
    @AustenThis is an excellent question. Why aren’t AI models just training themselves already? They theoretically can, and kind of are, but they don’t have the data/evals/gyms required to do so. A short reading list: Bottleneck is the environment, not compute https://medium.com/@shuchaobi/ais-next-bottleneck-isn-t-compute-it-s-the-environment-0eaec31888f5 RL envs cost real money ($20k–$300k/env, and this is for simulated ones which are just kinda crappy IMO) https://epoch.ai/gradient-updates/state-of-rl-envs We’re running out of human text https://epoch.ai/publications/will-we-run-out-of-data-limits-of-llm-scaling-based-on-human-generated-data Eval is the bottleneck https://ysymyth.github.io/The-Second-Half/ Verifier’s law https://www.jasonwei.net/blog/asymmetry-of-verification-and-verifiers-law What labs buy: Foody on RL envs https://www.youtube.com/watch?v=a00xIn5kwhM Economy as RL environment machine https://www.mercor.com/blog/the-economy-will-become-an-rl-environment-machine/ APEX-Agents generalization https://www.mercor.com/blog/generalization-results-from-training-on-the-apex-agents-dev-set/ Etna: ~$1B/yr on external data, supply-constrained Surge Tuesday (can it get through a workday?) https://surgehq.ai/blog/tuesday-frontier-work-index Dario: task/process distribution, not more web text https://www.dwarkesh.com/p/dario-amodei-2 OSWorld 2.0 (~20% on long workflows) https://osworld-v2.xlang.ai/ https://arxiv.org/abs/2606.29537 SWE-bench = ticket + repo + tests https://www.swebench.com Karpathy: sucking supervision through a straw https://www.dwarkesh.com/p/andrej-karpathy Ilya: peak data / one internet https://www.reuters.com/technology/artificial-intelligence/ai-with-reasoning-power-will-be-less-predictable-ilya-sutskever-says-2024-12-14/