• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Changing the code around the same AI model can reportedly create up to a 6× performance gap

    The post describes Meta-Harness, a system that automatically improves the “harness” code around an AI model by giving an optimizing agent access to prior code, logs and execution traces.

    RP
    2 Sources, 18d ago, first seen 18d ago

    TLDR

    A post describing a Stanford and MIT paper says an AI model’s “harness”—the code controlling storage, retrieval, what the model sees and how workflows run—can substantially affect performance. It describes gaps of up to 6× with the same model on the same benchmark when the harness changes. The post says Meta-Harness automatically improves that code by giving an optimizing agent access to previous code, logs and execution traces. It reports a 7.7-point improvement on online text classification over a strong state-of-the-art context-management approach while using 4× fewer context tokens, plus an average 4.7-point gain on retrieval-augmented math reasoning across five held-out models on 200 International Mathematical Olympiad–level problems.

    Combined views

    16.9K

    2 Sources, first seen 18d ago

    Combined views

    16.9K

    2 Sources, first seen 18d ago

    321 likes
    321 likes
    36 comments
    324 saves
    100 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    36 comments
    324 saves
    100 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    2 Sources

    @rohanpaul_aiStanford + MIT paper on Model Harnesses shows that AI performance depends not just on the model itself, but on the surrounding system code — the “harness”. This is what decides what to store, retrieve, show to the model, and how the workflow runs. With the same underlying LLM, changing the harness can create up to a 6× performance gap on the same benchmark. They conclude the harness around a model matters as much as the model itself. The paper introduces Meta-Harness, an outer-loop system that automatically improves harness code. Instead of giving the optimizing agent only a score or a short summary of past attempts, it gives the agent rich access to prior code, logs, and execution traces through a filesystem-like setup. The idea is that better diagnostic visibility lets the system improve the harness more intelligently. What it found is pretty significant. - On online text classification, a 7.7-point improvement over a strong SOTA context management, while using 4X fewer context tokens. - On retrieval-augmented math reasoning, an average gain of 4.7 points across five held-out models on 200 IMO-level problems. On agentic coding, the discovered harnesses beat strong hand-engineered baselines on TerminalBench-2. The paper shifts attention from “which model is best?” to “how is the whole AI system designed?” For real deployments, harness design affects reliability, tool usage, context management, and failure recovery. ---- Paper – arxiv. org/abs/2603.28052 Paper Title: "Meta-Harness: End-to-End Optimization of Model Harnesses"

    2 Sources

    @rohanpaul_aiStanford + MIT paper on Model Harnesses shows that AI performance depends not just on the model itself, but on the surrounding system code — the “harness”. This is what decides what to store, retrieve, show to the model, and how the workflow runs. With the same underlying LLM, changing the harness can create up to a 6× performance gap on the same benchmark. They conclude the harness around a model matters as much as the model itself. The paper introduces Meta-Harness, an outer-loop system that automatically improves harness code. Instead of giving the optimizing agent only a score or a short summary of past attempts, it gives the agent rich access to prior code, logs, and execution traces through a filesystem-like setup. The idea is that better diagnostic visibility lets the system improve the harness more intelligently. What it found is pretty significant. - On online text classification, a 7.7-point improvement over a strong SOTA context management, while using 4X fewer context tokens. - On retrieval-augmented math reasoning, an average gain of 4.7 points across five held-out models on 200 IMO-level problems. On agentic coding, the discovered harnesses beat strong hand-engineered baselines on TerminalBench-2. The paper shifts attention from “which model is best?” to “how is the whole AI system designed?” For real deployments, harness design affects reliability, tool usage, context management, and failure recovery. ---- Paper – arxiv. org/abs/2603.28052 Paper Title: "Meta-Harness: End-to-End Optimization of Model Harnesses"