Report
A coding agent with shell access reportedly matched or beat four multi-agent ML systems
A post describing an Apple paper says the systems shared one codebase, model, hardware setup and 24-hour budget.
TLDR
A post describing an Apple paper says a well-prompted coding agent with shell and file access matched or beat four multi-agent ML systems. It says the systems were rerun on one codebase with the same model, hardware and 24-hour budget. With GLM 5.2, the minimal agent reportedly earned medals on 62.5% of Kaggle tasks, versus 47.1% for the best published harness. The post also says adding parallel agents and a message channel reduced its medal rate from 55.7% to 33.3%.
Combined views
1.7K
1 Source, first seen ago
