Announcement
SEAR benchmark tests how AI agents improve robot-control strategies
SEARโs creators say they tested seven models on 330 simulated tasks, two physics engines and real robots.
TLDR
The team behind SEAR says its framework evaluates how language-model agents improve robot-control policies through repeated runs and feedback. They say even text-only models such as DeepSeek could control robots. In their tests, Astra did best when watching and correcting every move live; Fable did best when building its own tools and policy in code first.
Combined views
5.4K
1 Source, first seen ago
