GPT-6 Luna (Max) joins Agent Arena's Pareto frontier
Agent Arena's later update puts the model at +1.59% net improvement and a $0.05 median cost per task, reversing an earlier post that said it had missed the frontier.
TLDR
Agent Arena initially ranked GPT-6 Luna (Max) at #23, with +1.6% net improvement across 8K real-world agentic sessions, and said it had missed its Pareto frontier. In a later September 28 update, Arena said Luna had joined the frontier at +1.59% net improvement and a $0.05 median cost per task. It said that cost was 94% lower than GPT-6 Sol (Max)'s and 98% lower than GPT-6 Astra (Max)'s.
Combined views
10K
1 Source, first seen 2h ago