• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI
    Announcement

    WeirdML v3 compares GPT 6.1 Sol, Sonnet 5.5 and Grok 4.7, with some results incomplete

    A post sharing the results says Sonnet 5.5 scores better than Opus 5, while Grok 4.7 is ahead of Kimi-K3.

    T(
    HI
    3 Sources, ,

    TLDR

    A post introduced WeirdML v3 as a fully agentic benchmark with 11 hand-made tasks involving unfamiliar data, unclear goals and limited feedback. In a later update, the same poster says GPT 6.1 Sol is very token-efficient, close to Astra but with a lower peak. Sonnet 5.5 scores better than Opus 5, and Grok 4.7 is ahead of Kimi-K3, though not all results are complete.

    Combined views

    11.6K

    3 Sources, first seen 8h ago

    Combined views

    11.6K

    3 Sources, first seen 8h ago

    172 likes
    8h ago
    first seen 8h ago
    172 likes
    13 comments
    28 saves
    11 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Featured Source
    13 comments
    28 saves
    11 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    3 Sources

    @htihleGPT 6.1 Sol, Claude Sonnet 5.5 and Grok 4.7 results on WeirdML v3. 6.1 Sol is very token efficient, close to Astra, but has a lower peak. Sonnet 5.5 scores better than Opus 5, and Grok 4.7 is ahead of Kimi-K3. Not all these results are complete, and more results are coming. I tested Deepseek 4.1 Flash in codex instead of opencode, but did not see a significant difference in performance, except for lower cost.8h
    @teortaxesTexRT @htihle: GPT 6.1 Sol, Claude Sonnet 5.5 and Grok 4.7 results on WeirdML v3. 6.1 Sol is very token efficient, close to Astra, but has a…5h

    3 Sources

    @htihleGPT 6.1 Sol, Claude Sonnet 5.5 and Grok 4.7 results on WeirdML v3. 6.1 Sol is very token efficient, close to Astra, but has a lower peak. Sonnet 5.5 scores better than Opus 5, and Grok 4.7 is ahead of Kimi-K3. Not all these results are complete, and more results are coming. I tested Deepseek 4.1 Flash in codex instead of opencode, but did not see a significant difference in performance, except for lower cost.8h
    @teortaxesTexRT @htihle: GPT 6.1 Sol, Claude Sonnet 5.5 and Grok 4.7 results on WeirdML v3. 6.1 Sol is very token efficient, close to Astra, but has a…5h