Report
AREX-2 is claimed to improve solutions with more test-time rounds
HuggingPapers describes the 27B agent as trained on verifiable machine-learning and algorithmic tasks.
TLDR
HuggingPapers says AREX-2 improves its solutions with more test-time rounds. It reports scores of 81.8 on MLE-bench Lite and 70.7 on Frontier-CS, and says the agent transfers to deep research.
Combined views
3.7K
2 Sources, first seen 8h ago
21 likes
