• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI
    Report

    A coding agent with shell access reportedly matched or beat four multi-agent ML systems

    A post describing an Apple paper says the systems shared one codebase, model, hardware setup and 24-hour budget.

    Rohan PaulRP
    1 Source, 1h ago, first seen 1h ago

    TLDR

    A post describing an Apple paper says a well-prompted coding agent with shell and file access matched or beat four multi-agent ML systems. It says the systems were rerun on one codebase with the same model, hardware and 24-hour budget. With GLM 5.2, the minimal agent reportedly earned medals on 62.5% of Kaggle tasks, versus 47.1% for the best published harness. The post also says adding parallel agents and a message channel reduced its medal rate from 55.7% to 33.3%.

    Combined views

    1.7K

    1 Source, first seen 1h ago

    Combined views

    1.7K

    1 Source, first seen 1h ago

    25 likes
    25 likes
    9 comments
    10 saves
    9 comments
    10 saves
    Featured Source

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    #16

    Today's Rank

    #16

    1 Source

    Rohan Paul@rohanpaul_aiNew Apple paper basically says progress in automated ML engineering has come from models and runtimes rather than the scaffolding built around them. It finds one well-prompted coding agent with shell and file access matched or beat 4 multi-agent ML systems, so skip the orchestration and start with a single session. Those systems add search trees, memory layers, and specialist agent teams. They were designed when a model could only write code, not run it. This paper re-ran every wrapper on 1 codebase with the same model, hardware, and 24-hour budget. Giving the model a shell instead of a chat box was the only change that clearly mattered. With GLM 5.2, the minimal agent medaled on 62.5% of Kaggle tasks against 47.1% for the best published harness. Adding parallel agents and a message channel actually dropped its medal rate from 55.7% to 33.3%. – arxiv. org/abs/2609.40303 Title: "How Much of a Harness Does a Strong Agent Need for Autonomous ML Engineering?"1h