WeirdML v3 debuts with 11 hand-made tasks for AI models
The announcement says models must produce results despite limited data, unspecified goals and/or very limited feedback.
TLDR
WeirdML v3 is described in its announcement as a fully agentic benchmark with 11 complex, hand-made tasks. Models must explore and understand unfamiliar data, develop machine-learning and data-analysis pipelines, and produce results under constraints that include limited data, unspecified goals and/or very limited feedback.
Combined views
28K
4 Sources, first seen 3h ago
WeirdML v3 debuts with 11 hand-made tasks for AI models
The announcement says models must produce results despite limited data, unspecified goals and/or very limited feedback.