• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Elie Bakouch Replies Zephyr Z9 Data Most Likely

    ML research engineer replies that the data is most likely from @zephyr_z9.

    ('
    KA
    T(
    17 Sources, 28d ago, first seen 28d ago

    TLDR

    Elie Bakouch, an ML research engineer training LLMs at Prime Intellect after earlier work at Hugging Face on SmolLM and FineWeb, posted a reply in an AI discussion. The message states simply that @zephyr_z9 data most likely. The post carries a research tag. The packet contains only this single reply and the author's self-description. No additional context, confirmation, or surrounding posts appear in the supplied lines.

    Combined views

    106.5K

    17 Sources, first seen 28d ago

    Combined views

    106.5K

    17 Sources, first seen 28d ago

    1.2K likes
    1.2K likes
    168 comments
    109 saves
    40 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    168 comments
    109 saves
    40 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    17 Sources

    @zephyr_z9@eliebakouch Why do u think this happened?? They changed the training data mix or is it a new attention arch
    @eliebakouch@zephyr_z9 data most likely
    @teortaxesTex> muse spark 1.3 doesn't seem to be better on other long context benchmarks => very likely benchmaxed on this specific task maaan, if the second-rate companies do to the AA what Google did to Arena, that sure would suck…
    @bindureddyMeta’s Muse Spark 1.3 only look good on benchmarks In real life, It just a GPT 5.6 Terra class model The benchmarks are being gamed severely
    @yacineMTB@teortaxesTex Everything is an rl env. If you want your task specifically to be done well, publish an RL env for it so it's overfit
    @yoavgoso: - any task/environment that you can score reliably and invest enough money (= an obscene amount) in bruteforcing its training, will be solved - every challenging and popular benchmark will meet this bar - abilities won't necessarily transfer how can we measure progress?

    17 Sources

    @zephyr_z9@eliebakouch Why do u think this happened?? They changed the training data mix or is it a new attention arch
    @eliebakouch@zephyr_z9 data most likely
    @teortaxesTex> muse spark 1.3 doesn't seem to be better on other long context benchmarks => very likely benchmaxed on this specific task maaan, if the second-rate companies do to the AA what Google did to Arena, that sure would suck…
    @bindureddyMeta’s Muse Spark 1.3 only look good on benchmarks In real life, It just a GPT 5.6 Terra class model The benchmarks are being gamed severely
    @yacineMTB@teortaxesTex Everything is an rl env. If you want your task specifically to be done well, publish an RL env for it so it's overfit
    @yoavgoso: - any task/environment that you can score reliably and invest enough money (= an obscene amount) in bruteforcing its training, will be solved - every challenging and popular benchmark will meet this bar - abilities won't necessarily transfer how can we measure progress?