Announcement
WeirdML v3 compares GPT 6.1 Sol, Sonnet 5.5 and Grok 4.7, with some results incomplete
A post sharing the results says Sonnet 5.5 scores better than Opus 5, while Grok 4.7 is ahead of Kimi-K3.
TLDR
A post introduced WeirdML v3 as a fully agentic benchmark with 11 hand-made tasks involving unfamiliar data, unclear goals and limited feedback. In a later update, the same poster says GPT 6.1 Sol is very token-efficient, close to Astra but with a lower peak. Sonnet 5.5 scores better than Opus 5, and Grok 4.7 is ahead of Kimi-K3, though not all results are complete.
Combined views
11.6K
3 Sources, first seen ago
