• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    WeirdML v3 debuts with 11 hand-made tasks for AI models

    The announcement says models must produce results despite limited data, unspecified goals and/or very limited feedback.

    Teortaxes▶️ (DeepSeek 推特🐋铁粉 2023 – ∞)T(
    elieEL
    Lisan al GaibLA
    5 Sources, ,

    TLDR

    WeirdML v3 is described in its announcement as a fully agentic benchmark with 11 complex, hand-made tasks. Models must explore and understand unfamiliar data, develop machine-learning and data-analysis pipelines, and produce results under constraints that include limited data, unspecified goals and/or very limited feedback.

    Combined views

    114.4K

    5 Sources, first seen 20d ago

    Combined views

    114.4K

    5 Sources, first seen 20d ago

    988 likes
    20d ago
    first seen 20d ago
    988 likes
    62 comments
    341 saves
    69 reposts
    62 comments
    341 saves
    69 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    5 Sources

    Håvard Ihle@htihleIntroducing WeirdML v3, a fully agentic benchmark featuring 11 complex hand-made tasks. Models must explore and understand unfamiliar data, develop ML and data analysis pipelines and produce results despite limited data, unspecified goals and/or very limited feedback. 1/820d
    Lisan al Gaib@scaling01RT @htihle: Introducing WeirdML v3, a fully agentic benchmark featuring 11 complex hand-made tasks. Models must explore and understand un…20d
    elie@eliebakouchprobably one of my favorite benchmarks, with very niche hand crafted ML tasks!20d
    Teortaxes▶️ (DeepSeek 推特🐋铁粉 2023 – ∞)@teortaxesTexExcellent work from Havard, and brutal results for open source. Make WeirdML hard again!… at least for non-frontier models. Though I have to note that V4.1 is the only model running in its non-native harness, off a non-native provider. One of those things can be changed…20d
    Jaime Sevilla@JsevillamolRT @htihle: Introducing WeirdML v3, a fully agentic benchmark featuring 11 complex hand-made tasks. Models must explore and understand un…19d

    5 Sources

    Håvard Ihle@htihleIntroducing WeirdML v3, a fully agentic benchmark featuring 11 complex hand-made tasks. Models must explore and understand unfamiliar data, develop ML and data analysis pipelines and produce results despite limited data, unspecified goals and/or very limited feedback. 1/820d
    Lisan al Gaib@scaling01RT @htihle: Introducing WeirdML v3, a fully agentic benchmark featuring 11 complex hand-made tasks. Models must explore and understand un…20d
    elie@eliebakouchprobably one of my favorite benchmarks, with very niche hand crafted ML tasks!20d
    Teortaxes▶️ (DeepSeek 推特🐋铁粉 2023 – ∞)@teortaxesTexExcellent work from Havard, and brutal results for open source. Make WeirdML hard again!… at least for non-frontier models. Though I have to note that V4.1 is the only model running in its non-native harness, off a non-native provider. One of those things can be changed…20d
    Jaime Sevilla@JsevillamolRT @htihle: Introducing WeirdML v3, a fully agentic benchmark featuring 11 complex hand-made tasks. Models must explore and understand un…19d