• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Y Combinator says a better harness lifts the same AI model from 30% to 95% on ARC-AGI

    YC says it gathered researchers and founders to explore AI harnesses—the software around models—and lessons from building an agent for every employee at the company.

    YC
    PI
    IR
    8 Sources, ,

    TLDR

    AI harnesses deserve to be treated as research, not dismissed as scaffolding or prompt engineering, Y Combinator argues. It says the same model weights that score 30% on ARC-AGI score 95% with a better harness. YC describes a discussion spanning self-improving harnesses, messaging between agents, personal AI on local devices and its own workplace harness, QM. It also says the discussion covers what YC learned building an agent for every employee.

    Combined views

    439.6K

    8 Sources, first seen 27d ago

    Combined views

    439.6K

    8 Sources, first seen 27d ago

    3.7K likes
    27d ago
    first seen 27d ago
    3.7K likes
    176 comments
    6.4K saves
    709 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    176 comments
    6.4K saves
    709 reposts
    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    8 Sources

    @yacinelearningPrime Agent: A Self-Improving RLM Harness with Seth Karten from Prime Intellect https://x.com/i/broadcasts/1pKkOXoWELjKj
    @ycombinatorHarnesses often get dismissed as just scaffolding, just prompt engineering, and not real research. But that couldn't be farther from the truth. The same model weights that score 30% on ARC-AGI score 95% with a better harness. So we gathered a group of researchers and founders working at the frontier to do a deep dive into the state of harnesses. We cover how we got to this point, the case for making your harness as expressive as possible, and what YC learned building an agent for every employee in the company. 00:00 - @FrancoisChauba1: Why harnesses matter 04:27 - Building an auto-researcher by accident 07:13 - A five minute history of harnesses 13:56 - Self-improving harnesses 18:35 - @sethkarten: Prime Agent, a self-improving RLM harness 21:50 - Context as an L1, L2, L3 cache 24:51 - From Turing machine to von Neumann computer 28:33 - Messaging between agents 30:04 - ARC-AGI results 33:09 - Emulator Bench and GPU kernels 37:30 - @JonSaadFalcon: OpenJarvis, personal AI on personal devices 38:26 - How far behind are local models 39:21 - The five primitives of a personal AI stack 42:47 - Letting cloud models optimize your local stack 43:53 - 800x cheaper than the cloud 45:58 - @josh__france and @jbellregan: QM, YC's agent harness for work 47:29 - A history of YC's internal agents 49:24 - OpenClaw and a fleet of 50 agents 51:04 - Pulling the brain out of the sandbox 54:43 - Letting the agent choose its own sandbox and model 57:16 - The grind tool: budgets on goals 58:50 - Agents don't understand social context
    @sethkartenMy YC Paper Club talk on Prime Agent is out. I talked about moving beyond naive prompting toward an agentic OS, and how eval-driven harness design can expose more of a model’s underlying capabilities through persistent computation, memory, and agent-to-agent communication.
    @irinarishRT @ycombinator: Harnesses often get dismissed as just scaffolding, just prompt engineering, and not real research. But that couldn't be fa…
    @PrimeIntellectRT @sethkarten: My YC Paper Club talk on Prime Agent is out. I talked about moving beyond naive prompting toward an agentic OS, and how ev…
    @weaviatepodcastEvery tool call an agent makes stacks another observation into one giant prompt, and models were never pretrained on sequences like that. Frontier labs spend enormous resources making those long trajectories in distribution. Recursive Language Models take the opposite path: hand the model its prompt as a variable and let it write code that spawns recursive calls over pieces of it, so each call sees a small, local problem it can solve. 🌀 Later in the episode Alex Zhang, the MIT PhD student behind RLMs, gets into PrimeAgent, the Prime Intellect harness whose only tool is a Python REPL, speculative execution for roughly 2x agent speedups, and pairing RLMs with ColBERT-style retrieval. 🎙️ Weaviate Podcast #142, in full: https://www.youtube.com/watch?v=iv0MtXS_DQo

    8 Sources

    @yacinelearningPrime Agent: A Self-Improving RLM Harness with Seth Karten from Prime Intellect https://x.com/i/broadcasts/1pKkOXoWELjKj
    @ycombinatorHarnesses often get dismissed as just scaffolding, just prompt engineering, and not real research. But that couldn't be farther from the truth. The same model weights that score 30% on ARC-AGI score 95% with a better harness. So we gathered a group of researchers and founders working at the frontier to do a deep dive into the state of harnesses. We cover how we got to this point, the case for making your harness as expressive as possible, and what YC learned building an agent for every employee in the company. 00:00 - @FrancoisChauba1: Why harnesses matter 04:27 - Building an auto-researcher by accident 07:13 - A five minute history of harnesses 13:56 - Self-improving harnesses 18:35 - @sethkarten: Prime Agent, a self-improving RLM harness 21:50 - Context as an L1, L2, L3 cache 24:51 - From Turing machine to von Neumann computer 28:33 - Messaging between agents 30:04 - ARC-AGI results 33:09 - Emulator Bench and GPU kernels 37:30 - @JonSaadFalcon: OpenJarvis, personal AI on personal devices 38:26 - How far behind are local models 39:21 - The five primitives of a personal AI stack 42:47 - Letting cloud models optimize your local stack 43:53 - 800x cheaper than the cloud 45:58 - @josh__france and @jbellregan: QM, YC's agent harness for work 47:29 - A history of YC's internal agents 49:24 - OpenClaw and a fleet of 50 agents 51:04 - Pulling the brain out of the sandbox 54:43 - Letting the agent choose its own sandbox and model 57:16 - The grind tool: budgets on goals 58:50 - Agents don't understand social context
    @sethkartenMy YC Paper Club talk on Prime Agent is out. I talked about moving beyond naive prompting toward an agentic OS, and how eval-driven harness design can expose more of a model’s underlying capabilities through persistent computation, memory, and agent-to-agent communication.
    @irinarishRT @ycombinator: Harnesses often get dismissed as just scaffolding, just prompt engineering, and not real research. But that couldn't be fa…
    @PrimeIntellectRT @sethkarten: My YC Paper Club talk on Prime Agent is out. I talked about moving beyond naive prompting toward an agentic OS, and how ev…
    @weaviatepodcastEvery tool call an agent makes stacks another observation into one giant prompt, and models were never pretrained on sequences like that. Frontier labs spend enormous resources making those long trajectories in distribution. Recursive Language Models take the opposite path: hand the model its prompt as a variable and let it write code that spawns recursive calls over pieces of it, so each call sees a small, local problem it can solve. 🌀 Later in the episode Alex Zhang, the MIT PhD student behind RLMs, gets into PrimeAgent, the Prime Intellect harness whose only tool is a Python REPL, speculative execution for roughly 2x agent speedups, and pairing RLMs with ColBERT-style retrieval. 🎙️ Weaviate Podcast #142, in full: https://www.youtube.com/watch?v=iv0MtXS_DQo